<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>LLMJury Blog</title>
  <subtitle>Run live A/B tests on your prompts and models, version prompts outside your codebase, and get a plain-English verdict backed by real statistics — LLM evaluation with proof, not gut feeling.</subtitle>
  <link href="https://llmjury.com/atom.xml" rel="self"/>
  <link href="https://llmjury.com/blog"/>
  <id>https://llmjury.com/blog</id>
  <updated>2026-07-16T12:00:00.000Z</updated>
  <author><name>LLMJury</name></author>
  <entry>
    <title>Offline evals aren’t enough: why prompts need testing on live traffic</title>
    <link href="https://llmjury.com/blog/offline-evals-are-not-enough"/>
    <id>https://llmjury.com/blog/offline-evals-are-not-enough</id>
    <updated>2026-07-16T12:00:00.000Z</updated>
    <published>2026-07-16T12:00:00.000Z</published>
    <summary>Golden datasets are unit tests, not proof. Distribution drift, eval overfitting, and the metrics that only exist in production — latency, cost, and what users do next.</summary>
    <category term="evaluation"/>
    <category term="LLM-as-judge"/>
    <category term="production"/>
  </entry>
  <entry>
    <title>Why prompt changes deserve A/B tests, not vibes</title>
    <link href="https://llmjury.com/blog/why-ab-test-prompts"/>
    <id>https://llmjury.com/blog/why-ab-test-prompts</id>
    <updated>2026-07-16T12:00:00.000Z</updated>
    <published>2026-07-16T12:00:00.000Z</published>
    <summary>Ten playground outputs can’t tell you what a prompt change does to real traffic. The case for measuring every prompt edit the way you measure every other production change.</summary>
    <category term="A/B testing"/>
    <category term="prompt engineering"/>
    <category term="experimentation"/>
  </entry>
  <entry>
    <title>Announcing LLMJury: statistically defensible A/B testing for LLM products</title>
    <link href="https://llmjury.com/blog/announcing-llmjury"/>
    <id>https://llmjury.com/blog/announcing-llmjury</id>
    <updated>2026-07-01T12:00:00.000Z</updated>
    <published>2026-07-01T12:00:00.000Z</published>
    <summary>Why “it feels better” is not an eval strategy, and how SRM gates, FDR correction, and an LLM judge make prompt experiments trustworthy.</summary>
    <category term="announcements"/>
    <category term="statistics"/>
    <category term="SRM"/>
    <category term="FDR"/>
  </entry>
</feed>
