The AI stack repriced from top to bottom
This week the AI stack repriced from top to bottom. Frontier models got cheaper, the fastest inference moved off the model and onto specialized chips, a leading coding tool was bought by a company that also sells models, and a European app builder raised at a valuation that treats natural-language software as a lasting market. Capability got cheaper while ownership got more concentrated.
Frontier prices fell twice in one week. xAI shipped Grok 4.6, matching GPT-5.6 Sol quality at roughly half the price, per VentureBeat on August 12. A day later Google shipped Gemini 3.7 Flash at $0.75 and $3.75 per million tokens for input and output (the unit models bill by, a little under a word each), per SiliconANGLE on August 13. When two labs reach the frontier and undercut it inside 48 hours, the win moves from raw capability to cost per finished task.
The model stopped being the slow part. OpenAI previewed an "Ultrafast" tier that runs GPT-5.6 Sol up to 14 times faster, around 750 tokens per second, on chips from Cerebras, per OpenAI on August 13. Once generation is no longer the bottleneck, the software that coordinates several model calls into one result (the orchestration layer) decides who wins.
A "neutral" tool picked a side. SpaceX completed a roughly $60 billion all-stock purchase of Anysphere, maker of the Cursor coding tool, per Market Business News on August 14. A coding assistant many developers adopted because it stayed even-handed across models now belongs to a company that also sells one.
Europe priced the category as permanent. Sweden's Lovable raised $400 million in fresh venture funding at a $13.3 billion valuation, backed by European investors, per TechRepublic on August 12. Valuing a natural-language app builder (software that turns a plain-words description into working code) at that level treats the category as a durable market rather than a weekend demo, and it puts a European name near the top of it.
Capable models kept getting smaller. Meta released Muse Glimmer, a 30-billion-parameter open-weight agentic model (its internals are free to download and run, and it can carry out multi-step tasks on its own), small enough to run on a single consumer graphics card, per SiliconANGLE on August 10. Local agents that carry no per-token fee are becoming a real option next to hosted services.
The week's sharpest contrary read sits under all of this, in how the buildout gets paid for. Nvidia convened a $500 billion financing alliance sold as lowering the risk of the AI buildout, and this week's analysis argues it mostly shifts that risk onto outside lenders while treating chips that lose value fast as if they were long-lived infrastructure. Read it in The $500 billion bet that chips age like real estate.
The two directions are pulling apart. Prices are falling and capable models are spreading to ordinary hardware, while the tools and the money wrapped around those models keep gathering into fewer hands. For an operator, the cheap part is now the model; the advantage, and the switching cost that comes with it, is moving to everything built around it. Who captures the value when AI gets cheap is a question we took up in Who Really Wins From Subsidized AI.