OpenAI Took the Case-Law Index and Thomson Reuters Plugged In
On September 17 OpenAI released Astra for Law, its GPT-6 Astra model tuned for legal work, attached to a search index of more than 230 million URLs of US case law and statutes. The index was built with the Free Law Project, a nonprofit publishing court opinions. Thomson Reuters, the legal profession's research incumbent, shipped a ChatGPT plugin that day, per OpenAI's announcement. That week Wayfair put a sales agent in the same app and TotalEnergies paid to build a model. The assumption that stopped holding: a curated index of public data is a product.
The position that was defensible until Thursday
Legal AI was built on retrieval. The case law is public, but a current, clean, cited index of it was expensive to assemble, and a wrong citation costs a lawyer enough that firms paid for the index and the editorial layer on top. That is the business of Westlaw, Thomson Reuters' research service and the default tool in most US firms, and it was the startup pitch too: rent the model, own the retrieval and the workflow. Harvey, the best-funded legal AI startup, raised $550 million on September 9 on that position, per its own announcement. Eight days later OpenAI made the retrieval a feature of its application programming interface, or API, the connection developers use to call a model. The announcement says the model will reach API customers including Harvey and a Swedish rival, Legora, and lists 26 partner plugins.
The strongest case that nothing broke is OpenAI's own. Its announcement says the index "complements the licensed content and specialist products firms rely on from providers such as Thomson Reuters," and its numbers are modest. The test set was Legal Research Bench, built by Vals AI, an independent evaluation firm. On 200 of its questions Astra for Law passed the full correctness check on 54%, against 38.7% for the underlying GPT-6 Astra model with web search alone. OpenAI ran the test itself. Nearly half the answers still fail. On this reading the raw index was never the product; the editorial summaries on each opinion, the citator that flags overruled cases, the workflow and the liability were, and those still belong to the incumbents. Thomson Reuters plugging in is then a distribution win.
That reading is fair, and it concedes the point. Everything it lists as still defensible sits above the index. The index itself, the thing the buyer's demo was about, is now a feature of the model vendor, and its coverage claim, more than 99.9% of published US precedential case law, is OpenAI's own and unaudited. The lawyer still reads every answer. The firm can now see a price for the search on its own, separate from everything above it.
Retail: the sale moved in front of the website
The second break came a day earlier, and this week's Newsletter, Anthropic asked to slow AI down, and OpenAI asked for $1.5 trillion, gave it one paragraph. On September 16 OpenAI began piloting Sponsored Agents, an ad format that, when tapped, opens a labeled conversation with the advertiser's own AI agent inside the app, per OpenAI. Wayfair, the US online furniture retailer, confirmed it is in the pilot. Angi, a US home-services marketplace, said a homeowner can describe a repair inside the sponsored chat and only then enter Angi's process for finding a contractor, per its September 16 statement.
The prior position was that the website is where the sale is made and the buyer's intent captured. A search ad bought a click to a product page, and the page did the selling. In the pilot the qualifying happens on OpenAI's surface, the brand's agent handles the selling, and the site receives a shopper who has already decided. Google made a related move the same day. It opened early access to a connection that lets outside agents, ChatGPT among them, control Nest thermostats and doorbells. Access is gated to a $20-a-month subscription tier for US customers, per TechCrunch. The platform hands rival agents its devices and charges for the access.
Oil: the buyer whose data nobody else has
On September 15 TotalEnergies announced a three-year joint program with Mistral, the French AI lab, worth more than €100 million in total, to build proprietary models for finding and developing oil and gas reservoirs, trained on the company's own subsurface data, per TotalEnergies. There is no nonprofit publishing seismic surveys. No lab can assemble a public index of what sits under TotalEnergies' licenses, so the company is paying to build models in a joint laboratory rather than renting one.
Data a lab can reproduce from public sources becomes a feature of the lab's model, at the price of an API call. Data captured in a conversation belongs to whoever owns the surface the conversation happens on. Data nobody else has is the one asset still worth building a model around, and this week one buyer committed more than €100 million to do it.
The evidence is thin in three places. The Thomson Reuters plugin is described only in OpenAI's release; its terms, and whether any Westlaw content flows through it, are unknown. The Sponsored Agents pilot is days old, US-only and, in Wayfair's own words, small scale, per its statement. TotalEnergies is one company. Three sectors moving the same way in one week is concurrence, and this piece claims nothing more.
Who should be uncomfortable
Law firms and legal departments paying a per-seat premium for a tool whose demo was retrieval. The question to put to the vendor before renewal is what remains once the index is a line in OpenAI's price list.
Startups whose proprietary dataset is public records, cleaned: court dockets, company filings, property records, government statistics. On September 17 the United Nations and Google launched a platform that lets AI agents pull official UN statistics directly, with a target of 80% of the UN system's datasets by 2027, per the two organizations' announcement. Cleaning those records is a service now, priced as one.
Agencies whose retainer is search-ad management and on-site conversion work for consumer brands. The qualifying conversation just moved to a surface the agency does not control and cannot measure.
The move
Sort every dataset you sell or pay for into three piles. Public: someone with more compute will index it, so stop charging for the index and charge for what you do above it. Platform-captured: if your agent will sell inside ChatGPT or answer inside Google's home products, negotiate now for the transcript and the handoff record, because the intent data lives on their side. Wayfair's pilot runs with guardrails on accuracy, transparency and handoffs to service, per its statement; the handoff is the part to ask for. Proprietary: if the data is genuinely yours and large, price the TotalEnergies option against the subscription before your vendor does.
Two things to watch. Whether the Thomson Reuters connector promised for later ships with its research content inside ChatGPT or stays a handoff. And the list price of Astra for Law when it opens to API customers, the number every legal AI vendor's margin will be measured against.
The position
For two years the pitch to a professional buyer was that the model is a commodity and the data is the moat. This week showed which data. A moat made of public records was always a head start, and a head start is a depreciating asset. We argued in Defensibility in the AI era that durable positions are built in a sequence, and the index was never the first step. What TotalEnergies bought is the honest version of a data moat: a model trained on what nobody else can get. Everyone else is selling a service, and should price it like one.