Site navigation

Model intelligenceLive feed · no published verdicts

Verdicts by the work, earned by the evidence.

HiveBase picks a preferred model per work class — and for coding, a preferred harness × model × role. A class publishes that pick only when the evidence we are allowed to show backs it; until then it shows method, not a model.

Tracked sources are named and linked. Artificial Analysis is internal-only. Sample sizes and confidence intervals are shown when we have them. There is no “#1.”

LAST REFRESHED
Oct 3, 2026, 06:30 AM UTC

The public Model Intel publication (slug current) is reachable, but no work class in it carries an evidence-backed verdict yet. Classes below show methodology only — no model, n, or interval is claimed. The ledger API returns this same snapshot.

Classes
11
Verdicts published
0
Auto-reverts recorded
12
Feed
live · empty
WORK CLASSES

Eleven jobs. No verdicts published yet.

Jump a class, or read them in order. Coding rows always name the harness. n is the first-party sample; intervals are shown when present. Classes without an evidence-backed verdict show methodology only — no model, n, or interval is claimed.

code.agentic

Agentic coding

Multi-file repository work: implement, recover, verify. Ranked as (harness, model, role).

n unpublished
Unpublished default · harness × model
No published default

Awaiting first publication for this class.

Currently serving · reverted
kimi-k2.7-code

Applicable: Multi-file repository work: implement, recover, verify. Ranked as (harness, model, role).

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • The same model in a different harness is a different row.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

code.frontend

Front-end / UI

Generated websites, components, and full-stack UI. Taste plus deployable quality.

n unpublished
Unpublished default · harness × model
No published default

Awaiting first publication for this class.

Currently serving
claude-opus-5-5

Applicable: Generated websites, components, and full-stack UI. Taste plus deployable quality.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • The same model in a different harness is a different row.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

code.review

Code review

Finding real defects without comment fatigue. Cross-family review is required.

n unpublished
Unpublished default · harness × model
No published default

Awaiting first publication for this class.

Applicable: Finding real defects without comment fatigue. Cross-family review is required.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • The same model in a different harness is a different row.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

reasoning.deep

Deep reasoning

Strategic, second-order, long-horizon analysis.

n unpublished
Unpublished default · model
No published default

Awaiting first publication for this class.

Currently serving
gpt-6-sol

Applicable: Strategic, second-order, long-horizon analysis.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

agentic.tools

Agentic tool use

Multi-step tool calling with schema fidelity and recovery.

n unpublished
Unpublished default · model
No published default

Awaiting first publication for this class.

Applicable: Multi-step tool calling with schema fidelity and recovery.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

writing.narrative

Long-form writing

User-visible prose, brand voice, and durable narratives.

n unpublished
Unpublished default · model
No published default

Awaiting first publication for this class.

Applicable: User-visible prose, brand voice, and durable narratives.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

extraction.structured

Structured extraction

Schema-valid JSON / Output.object without silent field loss.

n unpublished
Unpublished default · model
No published default

Awaiting first publication for this class.

Currently serving · reverted
gpt-oss-120b

Applicable: Schema-valid JSON / Output.object without silent field loss.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

longctx

Long context

Documents, transcripts, and large codebases. Effective window, not marketed size.

n unpublished
Unpublished default · model
No published default

Awaiting first publication for this class.

Currently serving · reverted
qwen3-6-plus

Applicable: Documents, transcripts, and large codebases. Effective window, not marketed size.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

vision.doc

Document / vision

PDF, OCR, diagram, and image understanding.

n unpublished
Unpublished default · model
No published default

Awaiting first publication for this class.

Currently serving · reverted
gemini-3.5-flash

Applicable: PDF, OCR, diagram, and image understanding.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

speed.classify

Classification / gates

Binary and cheap structured gates. Recoverable if wrong.

n unpublished
Unpublished default · model
No published default

Awaiting first publication for this class.

Applicable: Binary and cheap structured gates. Recoverable if wrong.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

intel.x

X / community intel

Native X search and community recency. Discovery, not deciding.

n unpublished
Unpublished default · model
No published default

Awaiting first publication for this class.

Applicable: Native X search and community recency. Discovery, not deciding.

  • Research-based defaults stay useful before HiveBase outcome data accumulates.
  • Outcome statistics are observational HiveBase-owned dogfood, not a causal ranking.
  • Artificial Analysis numbers are tracked internally and never rehosted.

No attributed evidence published for this class yet.

METHODOLOGY

How a verdict is allowed to exist.

Work classes are the spine. External benchmarks, community consensus, and HiveBase first-party outcomes all hang off the same eleven jobs so they stay commensurable. Fusion never promotes a default from community recency alone — that channel is capped and decays.

Coding is harness-keyed. The ranked unit is (harness, model, role). An executor on Claude Code is not comparable to the same weights inside Cursor or Codex. Review must be cross-family.

Outcomes stay separate. Research-based defaults can exist before HiveBase outcome data accumulates. Dogfood statistics appear in their own block with population, dates, denominator, and uncertainty. n < 20 is below the reporting threshold, not a causal override. Durable selection guidance lives in /docs/guides/model-intelligence. The same snapshot is at GET /api/models/ledger.

Artificial Analysis is not rehosted. AA indexes inform internal routing. This page will only say they are tracked and point at artificialanalysis.ai. Attributed sources (Terminal-Bench, SWE-Rebench, Design Arena, LMArena, BFCL, Epoch) appear as name + outbound link.

Saturated suites (SWE-bench Verified, HumanEval, retired open LLM leaderboards) are weight-zero. A canary that fails the quality floor auto-reverts; those events are public when the live feed includes them.

CHANGELOG

What moved — including auto-reverts.

Promotions that did not hold stay on the record. Auto-reverts are not hidden behind a successful later default.

  1. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: claude-opus-4-8-1m → claude-opus-4-8-1m

  2. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  3. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: kimi-k2.6 → kimi-k2.6

  4. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  5. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on vision.doc: gemini-3.1-pro → gemini-3.1-pro

  6. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  7. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on extraction.structured: gpt-oss-120b → gpt-oss-120b

  8. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: qwen3-6-plus → qwen3-6-plus

  9. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  10. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  11. Sep 7, 2026, 07:16 PM UTCauto revert

    Rollback on reasoning.deep: gpt-5.6-sol → gpt-5.6-sol

  12. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: claude-opus-4-8-1m → claude-opus-4-8-1m

  13. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  14. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.6 → kimi-k2.6

  15. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  16. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on vision.doc: gemini-3.1-pro → gemini-3.1-pro

  17. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  18. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on extraction.structured: gpt-oss-120b → gpt-oss-120b

  19. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: qwen3-6-plus → qwen3-6-plus

  20. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  21. Sep 7, 2026, 08:00 AM UTCauto revert

    Rollback on longctx: qwen3-6-plus → qwen3-6-plus

  22. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  23. Sep 7, 2026, 08:00 AM UTCpromotion

    Routing change on reasoning.deep: gpt-5.6-sol → gpt-5.6-sol

  24. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: claude-opus-4-8-1m → claude-opus-4-8-1m

  25. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  26. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.6 → kimi-k2.6

  27. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  28. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on vision.doc: gemini-3.1-pro → gemini-3.1-pro

  29. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on vision.doc: gemini-3.5-flash → gemini-3.5-flash

  30. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on extraction.structured: gpt-oss-120b → gpt-oss-120b

  31. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: qwen3-6-plus → qwen3-6-plus

  32. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  33. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on longctx: qwen3-6-plus → qwen3-6-plus

  34. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on code.agentic: kimi-k2.7-code → kimi-k2.7-code

  35. Sep 7, 2026, 06:01 AM UTCpromotion

    Routing change on reasoning.deep: gpt-5.6-sol → gpt-5.6-sol

SOURCE RECEIPT

Last refresh by source.

Ingest health for this snapshot. A skipped or empty refresh is shown, not implied successful. Artificial Analysis appears here as a tracked source, never as a copied score.

Last refresh time and status for each Model Intel source
SourceFinishedStatus
research_pollOct 3, 2026, 06:25 AM UTCok
research_pollOct 3, 2026, 06:20 AM UTCok
research_pollOct 3, 2026, 06:15 AM UTCok
research_pollOct 3, 2026, 06:10 AM UTCok
research_pollOct 3, 2026, 06:05 AM UTCok
research_pollOct 3, 2026, 06:00 AM UTCok
fusionOct 3, 2026, 06:00 AM UTCok
overlayOct 3, 2026, 06:00 AM UTCok
research_pollOct 3, 2026, 05:55 AM UTCok
research_pollOct 3, 2026, 05:50 AM UTCok
research_pollOct 3, 2026, 05:45 AM UTCok
research_pollOct 3, 2026, 05:40 AM UTCok
research_pollOct 3, 2026, 05:35 AM UTCok
research_pollOct 3, 2026, 05:30 AM UTCok
research_pollOct 3, 2026, 05:25 AM UTCok
research_pollOct 3, 2026, 05:20 AM UTCok
research_pollOct 3, 2026, 05:15 AM UTCok
research_pollOct 3, 2026, 05:10 AM UTCok
research_pollOct 3, 2026, 05:05 AM UTCok
first_partyOct 3, 2026, 05:00 AM UTCok
swe_rebenchOct 3, 2026, 05:00 AM UTCskipped
hf_hubOct 3, 2026, 05:00 AM UTCskipped
wulong_arenaOct 3, 2026, 05:00 AM UTCok
design_arenaOct 3, 2026, 05:00 AM UTCskipped
epoch_eciOct 3, 2026, 05:00 AM UTCerror
bfclOct 3, 2026, 05:00 AM UTCok
research_pollOct 3, 2026, 05:00 AM UTCok
terminal_benchOct 3, 2026, 05:00 AM UTCok
openrouter_dataOct 3, 2026, 05:00 AM UTCok
openrouterOct 3, 2026, 05:00 AM UTCerror
lmarena_textOct 3, 2026, 05:00 AM UTCerror
artificial_analysisTracked internally — artificialanalysis.aiOct 3, 2026, 05:00 AM UTCerror
overlayOct 3, 2026, 05:00 AM UTCok
research_pollOct 3, 2026, 04:55 AM UTCok
research_pollOct 3, 2026, 04:50 AM UTCok
research_pollOct 3, 2026, 04:45 AM UTCok
research_pollOct 3, 2026, 04:40 AM UTCok
research_pollOct 3, 2026, 04:35 AM UTCok
research_pollOct 3, 2026, 04:30 AM UTCok
research_pollOct 3, 2026, 04:30 AM UTCok
research_pollOct 3, 2026, 04:25 AM UTCok
research_pollOct 3, 2026, 04:20 AM UTCok
research_pollOct 3, 2026, 04:15 AM UTCok
research_pollOct 3, 2026, 04:10 AM UTCok
research_pollOct 3, 2026, 04:05 AM UTCok
research_pollOct 3, 2026, 03:55 AM UTCok
research_pollOct 3, 2026, 03:50 AM UTCok
research_pollOct 3, 2026, 03:45 AM UTCok
research_pollOct 3, 2026, 03:40 AM UTCok
research_pollOct 3, 2026, 03:35 AM UTCok
research_pollOct 3, 2026, 03:30 AM UTCok
research_pollOct 3, 2026, 03:25 AM UTCok
research_pollOct 3, 2026, 03:20 AM UTCok
research_pollOct 3, 2026, 03:15 AM UTCok
research_pollOct 3, 2026, 03:10 AM UTCok
research_pollOct 3, 2026, 03:05 AM UTCok
overlayOct 3, 2026, 03:00 AM UTCok
research_pollOct 3, 2026, 03:00 AM UTCok
research_pollOct 3, 2026, 02:55 AM UTCok
research_pollOct 3, 2026, 02:50 AM UTCok
research_pollOct 3, 2026, 02:45 AM UTCok
research_pollOct 3, 2026, 02:40 AM UTCok
research_pollOct 3, 2026, 02:35 AM UTCok
research_pollOct 3, 2026, 02:30 AM UTCok
research_pollOct 3, 2026, 02:25 AM UTCok
research_pollOct 3, 2026, 02:20 AM UTCok
research_pollOct 3, 2026, 02:15 AM UTCok
research_pollOct 3, 2026, 02:10 AM UTCok
research_pollOct 3, 2026, 02:05 AM UTCok
research_pollOct 3, 2026, 02:00 AM UTCok
overlayOct 3, 2026, 02:00 AM UTCok
research_pollOct 3, 2026, 01:55 AM UTCok
research_pollOct 3, 2026, 01:50 AM UTCok
research_pollOct 3, 2026, 01:45 AM UTCok
research_pollOct 3, 2026, 01:40 AM UTCok
research_pollOct 3, 2026, 01:35 AM UTCok
research_pollOct 3, 2026, 01:30 AM UTCok
research_pollOct 3, 2026, 01:25 AM UTCok
research_pollOct 3, 2026, 01:20 AM UTCok
research_pollOct 3, 2026, 01:15 AM UTCok
research_pollOct 3, 2026, 01:10 AM UTCok
research_pollOct 3, 2026, 01:05 AM UTCok
research_pollOct 3, 2026, 12:55 AM UTCok
research_pollOct 3, 2026, 12:50 AM UTCok
research_pollOct 3, 2026, 12:45 AM UTCok
research_pollOct 3, 2026, 12:40 AM UTCok
research_pollOct 3, 2026, 12:35 AM UTCok
research_pollOct 3, 2026, 12:30 AM UTCok
research_pollOct 3, 2026, 12:25 AM UTCok
research_pollOct 3, 2026, 12:20 AM UTCok
research_pollOct 3, 2026, 12:15 AM UTCok
research_pollOct 3, 2026, 12:10 AM UTCok
research_pollOct 3, 2026, 12:05 AM UTCok
overlayOct 3, 2026, 12:00 AM UTCok
research_pollOct 3, 2026, 12:00 AM UTCok
research_pollOct 2, 2026, 11:55 PM UTCok
research_pollOct 2, 2026, 11:50 PM UTCok
research_pollOct 2, 2026, 11:45 PM UTCok
research_pollOct 2, 2026, 11:40 PM UTCok
research_pollOct 2, 2026, 11:35 PM UTCok
research_pollOct 2, 2026, 11:30 PM UTCok
research_pollOct 2, 2026, 11:25 PM UTCok
research_pollOct 2, 2026, 11:20 PM UTCok
research_pollOct 2, 2026, 11:15 PM UTCok
research_pollOct 2, 2026, 11:10 PM UTCok
research_pollOct 2, 2026, 11:05 PM UTCok
research_pollOct 2, 2026, 11:00 PM UTCok
overlayOct 2, 2026, 11:00 PM UTCok
research_pollOct 2, 2026, 10:55 PM UTCok
research_pollOct 2, 2026, 10:50 PM UTCok
research_pollOct 2, 2026, 10:45 PM UTCok
research_pollOct 2, 2026, 10:40 PM UTCok
research_pollOct 2, 2026, 10:35 PM UTCok
research_pollOct 2, 2026, 10:30 PM UTCok
research_pollOct 2, 2026, 10:25 PM UTCok
research_pollOct 2, 2026, 10:20 PM UTCok
research_pollOct 2, 2026, 10:15 PM UTCok
research_pollOct 2, 2026, 10:10 PM UTCok
research_pollOct 2, 2026, 10:05 PM UTCok
research_pollOct 2, 2026, 10:00 PM UTCok
overlayOct 2, 2026, 10:00 PM UTCok
research_pollOct 2, 2026, 09:55 PM UTCok
research_pollOct 2, 2026, 09:50 PM UTCok
research_pollOct 2, 2026, 09:45 PM UTCok
research_pollOct 2, 2026, 09:40 PM UTCok
research_pollOct 2, 2026, 09:35 PM UTCok
research_pollOct 2, 2026, 09:25 PM UTCok
research_pollOct 2, 2026, 09:20 PM UTCok
research_pollOct 2, 2026, 09:15 PM UTCok
research_pollOct 2, 2026, 09:10 PM UTCok
research_pollOct 2, 2026, 09:05 PM UTCok
research_pollOct 2, 2026, 09:00 PM UTCok
overlayOct 2, 2026, 09:00 PM UTCok
research_pollOct 2, 2026, 08:55 PM UTCok
research_pollOct 2, 2026, 08:50 PM UTCok
research_pollOct 2, 2026, 08:45 PM UTCok
research_pollOct 2, 2026, 08:40 PM UTCok
research_pollOct 2, 2026, 08:35 PM UTCok
research_pollOct 2, 2026, 08:30 PM UTCok
research_pollOct 2, 2026, 08:25 PM UTCok
research_pollOct 2, 2026, 08:20 PM UTCok
research_pollOct 2, 2026, 08:15 PM UTCok
research_pollOct 2, 2026, 08:10 PM UTCok
research_pollOct 2, 2026, 08:05 PM UTCok
overlayOct 2, 2026, 08:00 PM UTCok
research_pollOct 2, 2026, 08:00 PM UTCok
research_pollOct 2, 2026, 07:55 PM UTCok
research_pollOct 2, 2026, 07:50 PM UTCok
research_pollOct 2, 2026, 07:45 PM UTCok
research_pollOct 2, 2026, 07:40 PM UTCok
research_pollOct 2, 2026, 07:35 PM UTCok
research_pollOct 2, 2026, 07:30 PM UTCok
research_pollOct 2, 2026, 07:25 PM UTCok
research_pollOct 2, 2026, 07:20 PM UTCok
research_pollOct 2, 2026, 07:15 PM UTCok
research_pollOct 2, 2026, 07:10 PM UTCok
research_pollOct 2, 2026, 07:05 PM UTCok
research_pollOct 2, 2026, 07:00 PM UTCok
overlayOct 2, 2026, 07:00 PM UTCok
research_pollOct 2, 2026, 06:55 PM UTCok
research_pollOct 2, 2026, 06:50 PM UTCok
research_pollOct 2, 2026, 06:45 PM UTCok
research_pollOct 2, 2026, 06:40 PM UTCok
research_pollOct 2, 2026, 06:35 PM UTCok
research_pollOct 2, 2026, 06:30 PM UTCok
research_pollOct 2, 2026, 06:25 PM UTCok
research_pollOct 2, 2026, 06:20 PM UTCok
research_pollOct 2, 2026, 06:15 PM UTCok
research_pollOct 2, 2026, 06:10 PM UTCok
research_pollOct 2, 2026, 06:05 PM UTCok
overlayOct 2, 2026, 06:00 PM UTCok
research_pollOct 2, 2026, 06:00 PM UTCok
research_pollOct 2, 2026, 05:55 PM UTCok
research_pollOct 2, 2026, 05:45 PM UTCok
research_pollOct 2, 2026, 05:40 PM UTCok
research_pollOct 2, 2026, 05:35 PM UTCok
research_pollOct 2, 2026, 05:30 PM UTCok
research_pollOct 2, 2026, 05:25 PM UTCok
research_pollOct 2, 2026, 05:20 PM UTCok
research_pollOct 2, 2026, 05:15 PM UTCok
research_pollOct 2, 2026, 05:10 PM UTCok
research_pollOct 2, 2026, 05:05 PM UTCok
research_pollOct 2, 2026, 04:55 PM UTCok
research_pollOct 2, 2026, 04:50 PM UTCok
research_pollOct 2, 2026, 04:45 PM UTCok
research_pollOct 2, 2026, 04:40 PM UTCok
research_pollOct 2, 2026, 04:35 PM UTCok
research_pollOct 2, 2026, 04:30 PM UTCok
research_pollOct 2, 2026, 04:25 PM UTCok
research_pollOct 2, 2026, 04:20 PM UTCok
research_pollOct 2, 2026, 04:15 PM UTCok
research_pollOct 2, 2026, 04:10 PM UTCok
research_pollOct 2, 2026, 04:05 PM UTCok
overlayOct 2, 2026, 04:00 PM UTCok
research_pollOct 2, 2026, 04:00 PM UTCok
research_pollOct 2, 2026, 03:55 PM UTCok
research_pollOct 2, 2026, 03:50 PM UTCok
research_pollOct 2, 2026, 03:45 PM UTCok
research_pollOct 2, 2026, 03:40 PM UTCok
research_pollOct 2, 2026, 03:35 PM UTCok

Frequently asked questions

Why not one ranking?

A model that is right for cheap classification is the wrong unit for multi-file coding. HiveBase evaluates a verdict per work class and publishes it only when attributed evidence backs it — before that, a class shows methodology, not a pick. For code.agentic, code.frontend, and code.review the unit is (harness, model, role) — the same model on Claude Code versus Cursor is not the same system.

Why are Artificial Analysis scores missing?

Artificial Analysis is tracked internally for routing. Their terms do not allow us to rehost numeric scores. When AA is a source, the page says so and links to artificialanalysis.ai.

What does an auto-revert mean?

If a canary drops below HiveBase's quality floor, routing returns to the previous default. Public auto-reverts are listed in the changelog so a promotion that did not hold is visible, not buried.

Is this live production data?

When the public snapshot is reachable and carries an evidence-backed verdict, the receipt reads live. A reachable publication with no published verdicts renders classes as unpublished — methodology only, not a measured dataset. If environment, table, or query fail, this page renders a labeled preview snapshot so the marketing site never depends on production being up.

Are first-party numbers a ranking?

No. HiveBase-owned dogfood outcomes appear separately with population, dates, denominator, and uncertainty. They are observational. Research-based defaults remain useful before that evidence accumulates. Existing eval consent does not authorize publishing customer traces.

Routing you can inspect.

HiveBase routes work along these classes; a verdict joins them only as evidence earns it. The same boundaries, receipts, and undo as everywhere else.