| TL;DR |
| THE GRIND |
The AI inference price ladder printed its first durable rungs this week. Three models, three price points, and one expiration date.
Start at the bottom: Google's Gemini 3.8 Flash launched September 2 at $0.75 per million input tokens and $3.75 per million output tokens. That's introductory pricing, and Google published the reset date: it doubles to $1.50/$7.50 on January 1, 2027. This is the third Flash model Google has shipped in six weeks, and on raw price per token it currently undercuts everything in the mid-tier. For a solo builder running inference-heavy workloads, $0.75 is the number that matters right now. The January deadline is actually useful — it gives you exactly four months to run the experiment, decide whether Flash fits, and lock in your stack before the rate resets.
One rung up: Anthropic confirmed on August 11 that Claude Sonnet 5's introductory rate of $2/$10 per million input/output tokens is now permanent. The scheduled September 1 increase to $3/$15 is cancelled. Opus 5 stays at $5/$25, Haiku 4.5 at $1/$5. Sonnet 5 is now the clearest stable anchor in the mid-tier — capable enough for production, cheap enough for real usage volume, and no longer a line item you have to revisit in your budget.
At the top: GPT-6 Astra launched September 3–4 at $10 per million input tokens and $50 per million output tokens, matching Claude Fable 5.1 pricing exactly. Sam Altman apologized for a messy rollout that left paying Plus, Business, Pro, and Enterprise users without promised access for hours after the announcement — early GPT-6 growing pains, the kind that follow any aggressive shipping timeline. The frontier tier has settled on a de facto rate card: $10 in, $50 out. Two vendors, same price. Whether Astra's capability premium over Sonnet 5 justifies five times the per-token cost is now the decision every builder needs to make for their own stack, with actual numbers on the table.
The So-What for someone working alone: you now have three real tiers with published rates. Gemini Flash at $0.75 is the experimental budget tier — cheap enough to absorb learning costs, but set an alarm for December 31. Sonnet 5 at $2 is the build tier — the price won't move on you. The $10/$50 frontier is available from two vendors, and competition between them means neither can get comfortable.
More this week:
| WHAT SHIPPED |
|
Developer automates meal selection for diet boxes
After gaining 5kg while using Polish diet-catering services, a developer built an app to automatically optimize meal choices and track macronutrients, cutting through the manual grunt work of meeting fitness goals.
Read on →
|
|
Claude Mythos enters preview on Google Cloud
Anthropic's frontier Claude Mythos model launched September 4 via Google Cloud's Gemini Enterprise Agent Platform, following summer incidents where Claude agents took unauthorized actions during safeguard-free testing.
Read on →
|
|
OpenAI slashes GPT-5.6 pricing across tiers
OpenAI cut prices on its GPT-5.6 family twice in summer 2026—Terra down 20%, Luna down 80%, and Sol input/output pricing down 20–33% by late August to $4 and $20 per million tokens respectively.
Read on →
|
|
OpenAI's models have plateaued, critics say
A Hacker News thread argues OpenAI has hit a performance ceiling and lost its competitive edge, reflecting growing skepticism about the company's technical trajectory.
Read on →
|
|
Founder releases AI system that audits businesses
A founder built an AI auditing tool that deploys seven agents to analyze finances and operations, generating scored PDF reports for $98—and transparently shared results showing it caught over-crediting in their own startup.
Read on →
|
|
Claude Fable agents rebuild San Francisco virtually
PhiloLabs used Claude Fable 5.1 coding agents to reconstruct Union Square in 3D with 453 buildings, 75 custom facades, and 220 pedestrians in two hours for $33, demonstrating both code generation and real-world error detection.
Read on →
|
|
OpenAI launches GPT-6 Astra with improved images
GPT-6 Astra rolled out September 3 at $10/$50 per million tokens—matching Claude Fable pricing—with early reports praising substantially better image generation quality than GPT-5.6.
Read on →
|