Skip to main content
Router evals are model scorecards. Create one from Evals in the dashboard or upload a CSV with the Dari CLI:
model_id and score must be the first two columns. Higher scores are better. thinking_level and notes are optional. The CLI also accepts an optional metadata_json column containing a JSON object. Model IDs must exactly match the provider-prefixed IDs enabled on the router, and each model/level pair must be unique.
Use --file - to read CSV data from standard input. The command validates the CSV before creating the scorecard and prints the created eval as JSON. Put a cost_per_task_usd number in a row’s metadata_json to record what the model cost per task. When two or more rows have one, the eval page plots score against cost per task and marks the Pareto frontier: the models no other model beats on both. dari eval list returns each eval’s picks, its cheapest and best-scoring rows. Public Evals are synced by Dari from public leaderboards. They include SelfBench releases: coding tasks rebuilt from real repositories’ merged pull requests, each scored with accuracy and cost per task. Filter scorecards by model from the filter button on Evals, or with the CLI. A scorecard matches when it scores every listed model. Add thinking levels after a colon to require a score at any of those levels for that model:
Import the eval from the router create or edit page. Dari prefers a score matching both model and reasoning level, then falls back to that model’s row with a blank level. Scores inform selection; they do not define a fixed formula.