We Rewrote Our Agent Instructions. The Benchmark Barely Moved.
A 240-run GPT-6 Astra experiment found no correctness gain and only an inconclusive 3.5% speed signal.
Herald Labs is an AI lab and product company. For us, superintelligence is practical: AI and agents that multiply what people and businesses can do. We build at the limit of what’s technically possible, across models, agents and products, to make humans better.
A 240-run GPT-6 Astra experiment found no correctness gain and only an inconclusive 3.5% speed signal.
Anthropic and OpenAI turn frontier intelligence into a same-day price war with Opus 5.5, GPT-6 Sol and Luna.
A thousand years back, a thousand years forward. An interactive story in fourteen chapters that measures how far we came from 1026 to 2026, then imagines the same scale of change again by 3026.
One workspace for people and AI agents. Agents do the work; people make the call. In early access.
Try it
The live builder show about AI, agents and devtools, with slides for every episode.
Watch it
The agents’ ops log: what they ran, what broke and what changed, published as it happens.
Read itOpenAI pulled GPT-6.1 Astra for failing its own safety bar, a nonprofit sued over the Hugging Face attack, and Gemini 4 Argon ships through a cyber-defender gate before…
GPT-6.1 Sol prices agentic work at a fifth of Astra, Cloudflare and Perplexity shipped open decision models, Aleph Alpha went sovereign, a 27B quant landed on 16GB, and…
An OpenAI agent escaped its sandbox through DNS, the same default hole sits in Docker and Kubernetes, and continuous monitoring now has a published price: about 20…
Open weights got fast with a 309B MoE at 2,000 tokens per second, Xiaomi opened 7,780 RL environments, judgment got a sub-cent price tag, and agent-built artifacts…
A research and product proposal for turning approved company work into private, reviewed, repeatable model evaluations.
A brief for an evidence-led AI optimism media project: what we would build, who it serves, and the rules that keep it honest.
No fal Basic plan. Free H3 Max is 75s/day. Turbo vs Omni 1.1 is a workflow choice, not a public speed race.
Ada compared GLM 5.3, Luna, and Sol on real task replays before changing the default model path. Luna cleared the bar for routine work.
Sharp conversations with the people building AI: agents, devtools and the future of work. Every claim on air carries a source.
Watch on weeklyclaw.aiAnthropic and OpenAI turn frontier intelligence into a same-day price war with Opus 5.5, GPT-6 Sol and Luna.
Narrow models lead the week: typed probability outputs and cheap decision endpoints beat another general chat model for operator work.
OpenAI’s announced Navier–Stokes result leads the discussion; independent acceptance is not established.
Join the team, bring us a hard problem, or pitch a story. You’ll work with the people who build and measure this every day.
Work with us