Methodology
effiq ranks model reasoning variants for token-efficiency: high capability per real or estimated task dollar. The default gate keeps Artificial Analysis Intelligence Index ≥ 40. Users can move that floor and reweight metrics.
Sources
- Artificial Analysis — intelligence, coding, agentic indexes; measured cost per task; throughput and TTFT.
- OpenRouter — live provider prices, context, endpoint latency/throughput percentiles.
- Cursor — subscription-pool variants, effort/fast/thinking axes, published token prices.
WhatLLM and LLM Stats scrapers are configured but skipped until robots/terms checks pass. Their derived scores never override AA or official provider prices.
Efficiency Score
Among variants that pass the intelligence floor:
- Percentile-rank intelligence, coding, agentic, and throughput.
- Log-normalize and invert task cost and latency.
- Combine with user weights (default 35/15/10/30/5/5).
- Apply evidence-coverage and approximation penalties.
Raw capability per dollar (domain score ÷ task cost) is shown separately so near-zero prices cannot silently dominate.
Usage profiles
Eight profiles change domain evidence mix, default weights, and workload token assumptions: General, Coding, Agents, Math & Science, Finance, Research, Writing & Literature, Multimodal.
Approximations
- Exact measured variant
- Same-family interpolation between efforts
- Nearest-effort extrapolation
- Family aggregate + effort curve
- Insufficient data (no fabricated value)
Conservative ranking uses the less favorable bound of an estimate (higher cost / lower capability) when confidence is limited.
Sync
npm run sync refreshes the matrix. On deployment, a GitHub Actions scheduled workflow runs daily at 04:00 UTC and refreshes the matrix automatically.