Skip to content
← All posts

LLM-as-judge

Using a model to score another model’s output: the biases it brings, the budget it needs, and why a consistent judge beats an accurate one for comparing two arms.

3 posts

Other subjects