Cameron Otsuka

Why AI Companies Are Hunting for Data Nobody Else Has

Metadata
  • Description: The next AI race is for data that hasn’t already been consumed.
  • Publication: Inference Draft 2026-34
  • Published:
  • Last Modified:
  • Type: newsletter
  • Tags: ai
  • POSSE: Substack 
Glowing Unique Data Cube

The first question you would probably ask if I told you that Google bought Spirit Airlines’ data at auction, including internal emails, documents, booking information, and other corporate data, is whether Google would use it to improve Google Flights (Google’s flight booking tool) or start its own airliner.

The answer, is no. As far as I’m aware, Google is not planning to launch Google Airlines anytime soon. Instead, the data provides Google with a large corpus of real, internal, corporate communication. LLMs have already been trained on countless “business communication” articles, management books, and email templates. The harder data to find is a complete record of the communications between employees of a company over time and the decisions that were enacted because of them.

When someone asks AI for help drafting an email, a model trained on this Spirit Airlines data could provide outputs that better consider second order effects: how particular phrases are interpreted by others, the type of response it will likely receive, and how it should be worded to optimally reach your stated goal.

A day earlier, 404 Media revealed an investigation they had conducted on Amazon’s purchases of rare books. From 404 Media:

We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

The point of this operation isn’t to destroy the books themselves, but in the process of scanning and OCRing the text, the book will be destroyed. Anthropic has a similar initiative, Project Panama. These texts provide new, human-generated data untouched by AI outputs. In the case of rare books, they may even predate the internet and therefore have text that has never been seen online or used to train a competitor’s model.

A commodity market is emerging where datasets are priced based on uniqueness, cleanliness, and provenance. If the 2010s were the decade of “Big Data,” then the 2030s may be the decade of “Unique Data.”


Mine Print Hash

At the end of August, G20 finance and central bank ministers will gather in Asheville, NC for what may be the start of discussions towards a Bretton Woods for the stablecoin dollar era. Matt Dines and I touch on trade routing around the Hormuz and Red Sea choke points, including China’s recent “Ice Silk Road,” and early attempts across the globe at moving capital onto new rails.


Open Threads

China participates more deeply in the global order:

US tries to build infrastructure: