A new reasoning model called BDH-CQ solves ARC-AGI-1 tasks for $0.0007 each. That is a tenth of a cent.
It has 150 million parameters. Frontier models have thousands of times more. This one is small enough to run on a laptop, and it scores 29.5% on a benchmark designed to resist memorization.
For comparison: GPT-5.6 Luna (Low) scores 34.2%. Slightly better. At 11 times the cost, after OpenAI already cut prices by 80% in July.
The team behind it is Pathway, an AI lab building a post-Transformer architecture. Their result is a new point on the ARC-AGI-1 cost-accuracy frontier, and I think it is the most interesting cheap number in AI right now.
Disclosure: Pathway sent me their materials ahead of coverage. No payment involved. The read below is mine.
How it works
No chain-of-thought, no scratchpad
Reasoning models today write out their thinking as text. Every step produces tokens, and every token costs money.
BDH-CQ does not write anything down. It learns the task from the examples in front of it, stores what it learned in a recurrent memory, and works on the answer through repeated passes in latent space. The thinking happens inside the model, not in text.
That is where the 1,000x cost difference comes from. The tokens were the bill.

Official Cost-Accuracy Chart - in months it will reach 100% accuracy
Łukasz Kaiser, co-author of the paper that introduced the Transformer, replicated the result himself: "Pathway shows that model architecture, not just scale, can drive the next leap in AI reasoning." He is also a Pathway investor, so read that with both facts in hand.
Pathway CEO Zuzanna Stamirowska puts the claim in one line: "The bottleneck was never intelligence. It was design."
Sources: Paper (arXiv) · Hugging Face page · Pathway announcement · Code on GitHub
WHAT IT CANNOT DO
The honest part, and why I trust it more for it
29.5% is not the record. Frontier models reach 85% and more, at dollars per task. This result wins on price, not accuracy. Keep both numbers in view.
The paper also ran controlled tests on its own model, and one failure is striking. Ask it to apply one color rule, it works. Ask it to combine two color rules on the same object, it fails 0 out of 72 times. It can apply a rule. It cannot yet hold two rules together.
Two more caveats. The core architecture details and training recipe are proprietary, so outsiders cannot rebuild it. And the $0.0007 is computed from their own hardware time, while competitor costs on the ARC Prize leaderboard come from API prices. Not the same measurement.
So: a real result, honestly reported, not yet independently rebuilt. That is the accurate description.
Source: paper, Section 6, plus the open review thread on the Hugging Face page
Three episodes ago …
I wrote that the bottleneck in AI is no longer intelligence but observability. Then: correct answers got cheap, judgment did not.
This is the third data point in the same direction. A 150M model landing within five points of a frontier system says the constraint is not raw capability. It is architecture and cost.
One trade to notice. Chain-of-thought was expensive, but you could read it. Latent reasoning is cheap, and there is nothing to read. We are making reasoning affordable and less inspectable in the same step.
Cheap enough to run everywhere. Harder to audit when it does.
Was this forwarded to you? Start your 7-day free trial →
I build AI agents and GenAI systems for companies. If you're working on something like this, just reply to this email. I read every response.
Subscribe to AI Insider to read the rest.
Become an AI Insider to get access to this post and other subscriber-only content.
Join AI InsiderAI Insider members get:
- ✅ Full access to 100% of all content.
- ✅ Exclusive DEMOs, reports, and other premium content.
- ✅ Ad-free experience.
