Skip to main content
Admin › Evaluation
The evaluation feature measures AI response quality in three ways: Feedback (manual) · Arena & Leaderboard · Auto-Evaluation.
  • By combining direct user feedback, blind model comparison, and LLM-based automatic scoring, you can build a systematic quality-management framework.
  • When a measurement surfaces a problem, use Tracing to step through how that request was processed and diagnose the root cause.

Evaluation Methods

Three methods measure quality, and Tracing diagnoses the root cause of any problem they surface.
  • Measurement tells you what is bad; tracing tells you why.

Detailed Guides

Jump to the page that fits your goal.

Feedback (Manual)

Collect, view, and manage likes, dislikes, and comments

Arena & Leaderboard

Blind comparison · Elo-based model ranking

Auto-Evaluation

Judge-LLM automatic scoring — enable, results, statistics, export

Tracing

Trace the cause of low evaluation scores