Appearance
Evaluators
Evaluator types live at /evaluations/types. Open them from Evaluation Engine via Evaluation Types. This is the live list — it is not a legacy redirect and not /evaluations?tab=evaluators.
Section title: Evaluation types. Enable built-in evaluators or create custom ones.
Built-in types
ARMS ships 11 built-in types. Hallucination, Bias, and Toxicity are enabled by default.
| Evaluation Type | Description | Default |
|---|---|---|
| Hallucination | Factual inaccuracies, contradictions, fabricated information | Enabled |
| Bias | Discriminatory patterns across gender, ethnicity, age, religion | Enabled |
| Toxicity | Harmful, offensive, threatening, or hateful language | Enabled |
| Relevance | How well the response addresses the prompt | Disabled |
| Coherence | Logical flow and internal consistency | Disabled |
| Faithfulness | Alignment with provided context or source material | Disabled |
| Safety | Jailbreak, prompt injection, unsafe content | Disabled |
| Instruction Following | Adherence to instructions and formatting | Disabled |
| Completeness | Whether all parts of the query are addressed | Disabled |
| Conciseness | Avoids unnecessary verbosity | Disabled |
| Sensitivity | PII, credentials, confidential data exposure | Disabled |
Click a card to open /evaluations/types/[id]. Create custom types at /evaluations/types/new.
Online vs offline
- Online — Auto Evaluation or Run Evaluation on a live trace.
- Offline — Programmatic evaluations via
elsai_arms.eval. - Human — Manual Feedback.