Live AI agent hackathon · Monday, September 14, 2026
Fight poster: two robot boxers squaring off, labelled Astro 6 and Fable 5.1, Monday Sept 14, Live
InventoryHero
1
Where we are

Two weeks that moved the whole board.

  • Sep 1Anthropic ships Claude Fable 5.1 and Mythos 5.1.
  • Sep 2Google ships Gemini 3.8 Flash.
  • Sep 3OpenAI ships GPT-6 Astra. "A generational leap."
  • Sep 8OpenAI publishes a checked proof on Navier-Stokes.
  • Sep 10OpenAI ships GPT Image 2.5, Flare and Sunburst.
InventoryHero
2
Section 1 · OpenAI's two weeks

GPT-6 Astra: OpenAI calls it a generational leap.

Computer use
Drives a browser and a desktop.
State of the art at navigating computers and the web.
Cyber
First model rated Critical.
Finds unknown flaws on its own. That is why the release slipped.
Price
$10 in, $50 out.
Per million tokens. Identical list price to Fable 5.1.
The catch
"Recurrent depth" hides its reasoning.
Safety researchers cannot watch all of its thinking.
InventoryHero
3
Section 1 · OpenAI's two weeks

Astra helped close a 200 year old math problem.

Navier-Stokes: the equations for how fluids move. A $1M Millennium Prize question.
10,000
agents running at once
88 hrs
to the proof, on an unreleased model
17 hrs
for Astra to write it in Lean so a computer could check every step
2nd
Millennium Prize problem ever resolved
Terence Tao, as reported: "a carcass of raw meat on the table." True, but no why. OpenAI is not claiming the prize.
InventoryHero
4
Section 1 · OpenAI's two weeks

GPT Image 2.5 owns product and branding images.

RankModelLabScore
1gpt-image-2.5-sunburstOpenAI1419
2gpt-image-2 (medium)OpenAI1389
3gpt-image-2.5-flareOpenAI1388
4mai-image-2.6Microsoft AI1336
8gemini-3.1-flash-image (nano-banana-2)Google1269
Arena.ai Text-to-Image, Product, Branding & Commercial Design. Sep 7, 2026. 1.9M votes, 78 models. Sunburst and Flare preliminary.
Sunburst is the quality model. Flare is the fast one, half the latency of GPT Image 2.
InventoryHero
5
Section 2 · Anthropic's answer

Claude Fable 5.1: same brain, built for long jobs.

Shipped Sep 1
Coding, knowledge work, long runs.
Fixes root causes instead of taking shortcuts.
Price
$10 in, $50 out.
Unchanged from Fable 5. Cache reads dropped 75% to $0.25.
Real cost
25% to 45% cheaper.
Typical work, and the top end is for autonomous agent runs.
Mythos 5.1
Same model, fewer guardrails.
Gated to trusted programs. Fable is what you get in Claude Code.
InventoryHero
6
Section 3 · Head to head

On real agent work, Fable leads by a nose.

RankModelNet improvementConfirmed successPraise vs complaint
1Claude Fable 5.1 (Max)13.90%23.70%30.47%
2GPT-6 Astra (Max)11.90%19.54%34.41%
3Claude Opus 5 (Max)11.09%12.92%19.03%
Arena.ai Agent Arena, Overall. Sep 13, 2026. 1.69M sessions, 43 models.
Rank spreads overlap (1 to 4 vs 1 to 6). Call it a statistical tie at the top.
InventoryHero
7
Section 3 · Head to head

Tied on smarts. Astra is cheaper per task.

Fable 5.1GPT-6 AstraOpus 5Gemini 3.8 Flash
Intelligence Index53535141
Speed, tokens per second656051278
Cost per task$7.63$3.26$5.86$1.24
artificialanalysis.ai/models, Highlights. Viewed Sep 14, 2026. Max effort configurations; Fable 5.1 with fallback.
Same list price. Fable spends more tokens thinking.
InventoryHero
8
Honorable mention

Gemini 3.8 Flash is the value play.

41
Intelligence Index. Frontier level a few months ago.
278
tokens per second. Fastest on the chart.
$1.24
per task. A quarter of Astra, a sixth of Fable.
$0.75
in, $3.75 out per million, through Dec 31.
artificialanalysis.ai and blog.google, Sep 2026.
InventoryHero
9
Close

Pick the tool by the job.

Images, listings, one-shot tasks
GPT-6 Astra + GPT Image 2.5
Long agent runs and code
Claude Fable 5.1
Volume, speed, cheap
Gemini 3.8 Flash
Now, live. Replays and the next session:
inventoryhero.ai/ai-agents