Monday, August 3, 2026 · Run 2 · 1:11 PM ET
Astra Cracks Decade-Old Math Problems, Anthropic Passes OpenAI, and Marathon Coding Agents Arrive
8 stories from 4 newsletters
TL;DR
- OpenAI's unreleased "Astra" model solved 10 decades-old math and computer science problems for about $2K in compute.
- Anthropic reportedly passed OpenAI in revenue growth and valuation as Claude Code gains enterprise traction.
- Alibaba's Qwen3.8-Max ships a coding agent that works unsupervised for 10+ days, priced at a fifth of Fable 5.
AI Tips1
Use a "Gauntlet Loop" to force AI builds past a real quality bar
- Define the goal and name a real-world example that sets the quality bar.
- Break the goal into independent parts and assign each to a specialist builder agent.
- Assign a separate critic with fresh context to compare each part against the benchmark, blind if possible.
- Loop each part until its critic confirms it beats the benchmark — never let a builder grade its own work.
Use a Gauntlet Loop to complete this project. GOAL: [Describe the finished result.] REAL-WORLD EQUIVALENT: [Name or attach an excellent existing example that establishes the quality bar.] Break the goal into independent parts. Assign each part to a specialist builder. For every part, assign a separate critic with fresh context. The critic must inspect the generated artifact itself and compare it directly against the real-world equivalent. Where possible, compare them side by side without telling the critic which one is the reference. The critic may pass the work only if the generated artifact is better than the real-world equivalent. Otherwise, it must identify the largest specific gap and return the work for another iteration. Continue looping on every part until all critics pass it. Do not let builders evaluate their own work.
Source: The Neuron
AI News3
An unreleased OpenAI model, "Astra," solved 10 long-standing open problems
In mathematics, quantum complexity, and theoretical computer science (documented in a 249-page paper), including a group-theory conjecture unresolved since 1999 and three Erdős problems untouched for a decade — for roughly $2,000 in compute. Anthropic researcher Levent Alpoge said he reproduced 5 of the 10 proofs within 24 hours using Fable on a generic prompt.
Why it mattersA concrete jump in what these models can independently prove, not just summarize — useful ammunition next time a client asks whether AI can do genuinely novel work.
Source: The Rundown AI, Superhuman AI
Anthropic reportedly passed OpenAI in revenue growth and valuation
As Claude Code gains enterprise traction, while investors scrutinize OpenAI's cash burn.
Why it mattersWorth tracking given how much of our own delivery runs on Claude — this is a signal the tooling bet is paying off, not just a horse-race headline.
Source: The Neuron
Alibaba released Qwen3.8-Max
A 2.4-trillion-parameter model whose coding agent reportedly works unsupervised for 10+ days on a single project. It goes open-source next week, priced at roughly a fifth of Fable 5's API rate.
Why it mattersAnother credible, cheaper option to weigh when scoping AI-dev tooling for cost-sensitive clients — worth a bench test once the weights drop.
Source: The Rundown AI, The Neuron
Beyond AI4
Japan's Shinkansen remains the world's benchmark for rail service
Turning a six-hour drive into two hours at up to 200 mph with an average delay of just 1.6 minutes. The 1980s split of the national railroad into six competing companies let them cut unprofitable routes and reinvest non-rail revenue (hotels, retail near stations) into service and staff training.
The takeawayA clean case study on how breaking up a monopoly and diversifying revenue can fund the operational excellence a single centralized org never got around to.
Conversation starterOur client's ops org is basically the pre-1980s Japanese railroad — what would splitting it into accountable units actually free up?
Source: Morning Brew (Transportation Brew)
A Waymo employee called the police on two teen riders
After mistaking Orbeez gel-pellet toy guns for real firearms during live camera monitoring. Waymo remotely disabled the car and told the passengers it was a mechanical issue to keep them seated until officers arrived; the teens were released to their parents without arrest.
The takeawayNo direct relevance to your work, general awareness only — but a vivid illustration of always-on remote monitoring going wrong in a customer-facing product.
Conversation starterDid you see the Waymo story where a monitoring employee mistook a toy gun for real and called the cops? Wild case study in remote-monitoring false positives.
Source: Morning Brew (Transportation Brew)
Air taxis are being positioned as the fastest way to the airport
The US approved eight eVTOL pilot programs across 26 states in March, and Joby Aviation has already flown demo flights between Manhattan and JFK targeting 10-minute trips versus an hour-plus by car; fares could start near $150 before dropping toward $25 by 2030 if the business scales.
The takeawayNo direct relevance to your work, general awareness only.
Conversation starterWould you actually pay $150 to skip airport traffic in a drone taxi, or does that only make sense once it's $25?
Source: Morning Brew (Transportation Brew)
Free airline wifi is becoming a loyalty-program funnel, not a perk
Airlines increasingly gate onboard wifi behind loyalty sign-ups, and loyalty programs are now major revenue lines — topping $6B each in 2024 for American and Delta. Meanwhile Delta and United are testing stripped-down perks on cheaper fares to push flyers toward premium tiers.
The takeawayA clean parallel for any client debating whether to bundle a free feature to drive account creation versus protecting it as a paid tier.
Conversation starterAirlines are using free wifi purely to get you into their loyalty database — are we doing the equivalent with any of our own "free" features?
Source: Morning Brew (Transportation Brew)
Summarized from The Rundown AI, Superhuman AI, The Neuron, Morning Brew (Aug 2–3)
