Stssicila

Market Prices

Coin Price 24h
BTC Bitcoin
$77,931.8 +0.52%
ETH Ethereum
$2,447.27 +0.68%
SOL Solana
$105.02 +0.50%
BNB BNB Chain
$691.2 +0.07%
XRP XRP Ledger
$1.39 +0.20%
DOGE Dogecoin
$0.0852 +0.37%
ADA Cardano
$0.2004 -0.99%
AVAX Avalanche
$7.31 +0.55%
DOT Polkadot
$0.8389 -0.98%
LINK Chainlink
$11.4 +0.06%

Fear & Greed

68

Greed

Market Sentiment

Event Calendar

{{年份}}
28
03
unlock Arbitrum Token Unlock

92 million ARB released

18
03
unlock Sui Token Unlock

Team and early investor shares released

12
05
halving BCH Halving

Block reward halving event

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

Altseason Index

41

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
1
Bitcoin
BTC
$77,931.8
1
Ethereum
ETH
$2,447.27
1
Solana
SOL
$105.02
1
BNB Chain
BNB
$691.2
1
XRP Ledger
XRP
$1.39
1
Dogecoin
DOGE
$0.0852
1
Cardano
ADA
$0.2004
1
Avalanche
AVAX
$7.31
1
Polkadot
DOT
$0.8389
1
Chainlink
LINK
$11.4

🐋 Whale Tracker

🔵
0x91cf...d649
6h ago
Stake
10,617 BNB
🔴
0x0438...9ae5
1d ago
Out
3,585 ETH
🔵
0xee40...ee42
2m ago
Stake
15,696 BNB

💡 Smart Money

0x1e44...9aee
Arbitrage Bot
+$0.5M
86%
0xf554...1bcc
Institutional Custody
+$0.2M
67%
0x5ad9...790b
Market Maker
+$3.0M
64%

🧮 Tools

All →

Burned Books, Buried Liability: The $15 Billion AI Copyright Trade Nobody Is Pricing

Blockchain | PowerPomp |

Here is a data point no dashboard will show you. Over the past two years, an estimated several million physical books have been purchased in bulk at liquidation prices, had their spines severed by industrial guillotines, been fed through high-speed scanners at 300 dpi, converted into text through optical character recognition, and then been physically destroyed. Not by a totalitarian state. By the most richly capitalized companies in the history of commerce. Contract-worker job boards carry listings for "book scanning technicians." Logistics invoices track pallets of remaindered paperbacks into unmarked warehouses. The books are not the product. They are the raw material. The product is training tokens for artificial intelligence models, and the originals are discarded because disposal is cheaper than storage, return, or resale. The "burning" is not a cultural statement. It is a cost line.

I have a professional reflex at this point. I spent the DeFi Summer of 2020 running a high-frequency arbitrage bot on Uniswap v2, monitoring liquidity pool imbalances across Curve and Balancer, executing micro-trades to capture spread inefficiencies. The strategy earned 120% APY for six months, and then a flash loan attack on an integrated protocol froze liquidity, forcing me to manually intervene and pull $30,000 to safety in under an hour. That experience installed a permanent circuit breaker in my thinking: when reported yield is dramatically out of line with the risk-free rate, the trade is not free. It is a premium for bearing a risk the market has not priced. When an actor pays real, physical costs to manufacture one side of a trade, the trade is never about the physical asset. It is about the conversion rate between two stores of value. Paper books are being converted into training tokens at a conversion rate the legal market has not yet priced. That is an arbitrage. Arbitrage is just patience wearing a math mask.

The uncomfortable truth is that this is not a new episode. It is the third act of a data-acquisition play that began two decades ago. Google's Books project started scanning the world's libraries in 2004. Fifteen million volumes. The engineering culture optimized for volume, deferred to the legal department, and trusted the courts to catch up at a speed courts never move. It took a full decade of litigation before the company was forced to settle, and by then the scans were embedded in the institutional memory of one of the most powerful organizations on Earth. The lesson the surrounding industry absorbed was not "license first." It was "volume first, settlement later." That lesson has metastasized into an industrial practice.

By 2023, Bloomberg had reported that Meta held internal discussions about acquiring Simon & Schuster, an entire publishing house, specifically to convert its catalog into training data. The Atlantic had exposed Books3, a corpus of roughly 183,000 books assembled from pirate repositories, quietly making its way into the training pipelines of open-weight architectures. The New York Times had filed suit against OpenAI and Microsoft. The Authors Guild had followed. Getty Images had launched copyright actions against Stability AI on two continents. None of that stopped the physical pipeline. The reason is not ideological. It is a spreadsheet, and I have spent enough years modeling risk-adjusted yield to know that nobody walks away from a spreadsheet showing a 50x discount because an attorney says the risk is unpriced. The attorney is exactly why the discount exists.

This is where my verification reflex kicks in. The original reporting on the book-burning practice is, to be honest, thin. No named company. No dated invoices. No photographed warehouse. In my early career, I allocated my entire semester fund into the Status Network SNT presale and refused to trust the whitepaper's commitments; I tracked on-chain distribution against public team wallets and found a 40% insider concentration before the broader market did. That experience taught me a discipline: a claim without verifiable data is a narrative, not a fact. The book-burning story is a narrative with strong industry backing but weak documentary proof. What matters is not whether the specific allegation survives the fact check. What matters is that the underlying mechanism is real, multiple independent data points confirm the pattern, and the industry structure guarantees the incentive.

Let me establish the technical route, because industry context matters. The pipeline is three steps: source, digitize, delete. Step one: acquire books in bulk at wholesale remainder prices, between one and five dollars per volume with freight factored in. Step two: guillotine the spine so pages feed flat through batch scanners, capture at high resolution, run the OCR stack, clean the output, and load the text into the training corpus. Step three: dispose of the physical object. The disposal step reads as vandalism, but it is the cheapest line item in the entire operation. Returning books costs freight and restocking labor. Reselling them costs listing time and warehouse real estate. Grinding them into pulp costs cents per unit. An operations manager with a flow chart chose the most efficient annihilation on the menu. The books did not lose to ideology. They lost to logistics.

This is a deliberate, board-approved cost-engineering decision. A data engineer understands exactly why books are worth the physical hassle. Web-scraped text is noisy, truncated by paywalls and robots protocols, and contaminated by autogenerated spam. A full-length book carries long-range semantic structure, coherent reasoning chains, narrative arcs that train a model to track causality across thousands of tokens. That structural density is precisely what separates a frontier model from a chatbot that stumbles over the second paragraph. The AI labs are not paying for paper. They are paying for the scarcest input in the industry: clean, long-horizon text with a factual density the open web cannot provide. And because that input is scarce, the acquisition has pushed them into physical logistics.

Burned Books, Buried Liability: The $15 Billion AI Copyright Trade Nobody Is Pricing

The books are also not all the same. Some portion of this pipeline is certainly public-domain texts, orphan works, and out-of-print titles where the practical copyright holder is unreachable. The reports treat all books as equally protected, which is sloppy. But the strategic core of the operation is clear: a frontier lab needs contiguous, long-horizon text, and the most efficient physical source of that text happens to be a used-book wholesaler.

The most important number is not the price of the books. It is the transaction-cost asymmetry. Licensing one book requires identifying the copyright holder, establishing chain of title, negotiating the scope of use, executing a contract, and maintaining compliance records. That legal workflow is worth hundreds to thousands of dollars per volume before royalty. Multiply that by one million volumes, and the licensed pipeline becomes a multi-billion-dollar, multi-year commitment with an army of external counsel. The physical pipeline has no such friction. A used-book wholesaler does not demand a copyright audit. A liquidation crate does not arrive with a chain-of-title ledger. The purchase of a physical object confers the right to own, to hold, and to destroy that object. Under copyright law, it does not confer the right to copy it, digitize it, or train a model on its text. The gap between what is conferred and what is claimed is the entire trade. The industry is monetizing the distance between a receipt and a license.

Now I tax the risk. This is the section of my analysis I call the Risk Tax, because yield is never free; it is compensation for bearing specific systemic risks, and if you cannot name the risks, you have not priced the yield. Assume one frontier lab has pushed five million physical books through this pipeline. Say the all-in cost is five dollars per volume including scanning, labor, OCR compute, and disposal. That is twenty-five million dollars. Compare that to a hypothetical licensed corpus. Industry data-licensing conversations for high-quality narrative text, the kind of text that builds long-horizon reasoning, settle in the double-digit to triple-digit dollars per volume before legal overhead buries the deal. The licensed path for a multi-million-volume corpus runs into the billions. The discount on the physical path is not 30 percent. It is one to two orders of magnitude. That is not a business model. It is a liability weapon aimed at the balance sheet.

Let me make the downside precise. US copyright law provides statutory damages of up to $150,000 per infringed work, plus attorney's fees, when infringement is willful. A consolidated class action covering one hundred thousand registered titles yields a headline exposure of fifteen billion dollars. This is not a fantasy number. It is statutory arithmetic. The structural detail that separates this trade from ordinary market risk is the shape of the payoff. Copyright liability does not mark-to-market gradually. It lingers at zero for years while the pipeline compounds, and then it gaps to a massive number in a single judicial decision. The position is convex. It pays off incrementally, as a constant flow of high-quality training text, and it loses catastrophically and discretely. In payoff structure this is identical to a yield farmer holding a governance token that trades sideways for a year and then gaps 90 percent during an exploit. The yield is steady. The tail is the entire position.

Burned Books, Buried Liability: The $15 Billion AI Copyright Trade Nobody Is Pricing

I have watched this movie in a different theater. In DeFi, protocols advertised double-digit yields against collateral that had never been audited. The market priced the yields. It refused to price the liabilities. When the collateral was a token and the token was a promise, the promise failed in a week. The Terra collapse in 2022 was the same lesson at market scale: an algorithmic stablecoin promising twenty percent on nothing but faith in a minting mechanism drew in billions, and the drawdown did not wait for a governance vote. Here the collateral is a legal argument: "ownership of the physical object conveys a right to use the text." That argument is performing the same rhetorical function that "decentralization" performs in a DAO charter. It sounds robust. It produces no cover when the judge reads the statute. DAOs are compliance shields, and so is a receipt for a crate of pulped paperbacks. The companies running this pipeline are not rebels. They are timing the legal system, and the legal system is not obligated to honor their schedule.

The window is closing, and the mechanics are visible. The European Union's AI Act, Article 53, requires providers of foundation models to publish a sufficiently detailed summary of the training data used, and the transparency obligations are being operationalized now. China's Interim Measures for Generative AI Services explicitly prohibit training data that infringes intellectual property rights. The United States has no omnibus statute, so it has a litigation wave instead: the New York Times against OpenAI and Microsoft, the Authors Guild against the same class of defendants, Getty Images against Stability AI in multiple jurisdictions, and a docket of writer class actions. I have repeated until hoarse that emotional narratives are noise and regulatory text is signal. The regulatory text of three major jurisdictions is converging on one direction: training-data provenance will be audited, and the source of the corpus will be disclosed. When that day arrives, a model trained on burned paperbacks carries a disclosure problem its competitors do not share.

There is also an open-model dimension that the headline stories ignore. Books3 did not flow only into closed frontier labs; it flowed into open-weight models that were then downloaded and fine-tuned across the entire ecosystem. That means the liability is not concentrated in a handful of corporate defendants. It is diffused across thousands of derivative models, many generated by hobbyists with no legal counsel. Diffusion of liability does not reduce the exposure. It fractures it into thousands of claim sources, making a coordinated settlement impossible and raising the odds of an aggressive judicial precedent.

The market has not priced any of this. That is the entire thesis of this piece. AI training data is the largest unaudited ledger in human history, and unlike the blockchain rails I work with every day, it is also unauditable. On-chain settlement gives us verifiable records of token movement. The AI industry has built a parallel economy where the most valuable asset on the books, the training corpus, is a black box. In 2025, I pivoted a portion of my portfolio into the AI-crypto convergence thesis, investing in Render and Fetch.ai and betting on decentralized compute as the scarce layer for AI training. I built custom dashboards to track GPU utilization rates and agent transaction volumes on-chain. The compute layer of the AI economy has real telemetry. The data layer has none. That asymmetry is the tell of a market that priced growth and ignored provenance. When provenance is forced into the open, and it will be, a subset of models will be revealed as structurally compromised. The recognition event is a liability event.

Now the contrarian read, because the comfortable narrative is rarely the productive one. The dominant framing says the AI companies are villains and the destroyed books are a pure cultural loss. I want to dismantle three components of that framing, because each contains a blind spot that matters for anyone trying to position for the next two years.

The cultural loss is not homogeneous. A significant fraction of the books flowing through liquidation channels were already destined for pulping. A mass-market paperback with a two percent sell-through rate does not reach a reader. It reaches a remainder table, and then a landfill. Disposal of unsold stock is an ordinary event in publishing logistics; the novelty here is the digitization that happens before disposal, not the disposal itself. The romantic image of the Library of Alexandria vanishing in flames has no statistical counterpart in a warehouse of unsold thrillers. Critically, no one has actually counted these books. The source reports provide no title list, no language breakdown, no purchase manifests, no documentation beyond the aggregate allegation. The absence of evidence is the most reliable fact in this entire affair. Nobody who can count has counted.

The digital-preservation angle cuts both ways. If these companies had scanned the books and deposited the digital copies into a public archive, the operation would read as a rescue mission for orphaned texts. Instead, the scans are locked inside a private training corpus with no public access, no library deposit, no verification mechanism. The information survives, but its access has been privatized. That is a subtler injury than fire. Burning destroys the object. Locking the scan converts a public knowledge asset into a private competitive moat. The first is a crime against a warehouse. The second is a crime against a commons. The authors are discovering what NFT artists discovered after royalty enforcement collapsed: the revenue mechanism is only as strong as the platform's willingness to enforce it, and the platform has every incentive to look the other way.

The truly counter-intuitive point is this: the largest loser in this affair is unlikely to be the publisher. It is the AI industry itself. Consider the downstream effect of an adverse ruling on a model trained on a massive illicit corpus. Enterprise customers, regulated industries, and government agencies that adopted the model face their own compliance exposure. They will not respond with moral outrage. They will respond with procurement risk management. They will quietly stop deploying the model. Adoption collapses not because anyone is angry, but because liability is contagious. I watched this in DeFi: liquidity evaporated from protocols within hours of a proof-of-reserves audit revealing a mismatch. Liquidity is the court. The same mechanics will hit model usage the moment a data-provenance audit lands. Strategy is the art of surviving your own leverage, and a frontier lab leveraged on a pirated corpus is leverage without a stop-loss.

Let me close with the signal list, because a trader is only as good as their watching list. Three signals matter, and I watch them the way I watch liquidity pools during a depeg.

A legitimate licensing rail. If a major publisher begins listing "AI training corpus" as a standard commercial SKU, with per-volume or per-token pricing and a machine-readable license, the arbitrage window starts to close. Institutions will pay the premium to avoid the tail. Whoever occupies that lane, a publisher consortium or a brokerage for authorized corpora, owns the infrastructure opportunity of the next five years. I have seen this pattern before. When the dark pool of unregulated yield closes, capital does not leave the yield market. It migrates to audited venues. Watch the balance sheets of the publishing conglomerates; if a Pearson or a RELX starts hiring data-licensing sales teams, they have read the same tea leaves I have.

A judicial ruling on statutory damages for training data. Not a settlement. A ruling. Settlements are the cost of doing business. A ruling creates a priced benchmark, and the moment a benchmark exists, every balance sheet in the industry re-marks latent liability to market. The re-marking event will be violent for the companies carrying the largest unlicensed corpora, and it will be a gift to the compliance layer that can measure the exposure.

Data-provenance audit infrastructure. Just as DeFi demanded transparent reserves after the collapses, AI will demand transparent sources after the rulings. The fingerprinting technology to scan a corpus against registered works exists today. The demand does not yet. When regulation and litigation arrive together, the firms that built the audit tools own the map. I treat this as the cleanest long trade in the entire AI stack, because it is the only position that appreciates exactly when the risk realizes.

The books are gone. The scans are locked. The liability is compounding at the speed of the scanner. During the Terra collapse, I moved $200,000 into USDC and staked ETH within hours and shorted the failing ecosystem's native tokens, converting a potential catastrophe into an $85,000 gain. The survival protocol has not changed: position for the event, not the narrative. Volatility is the tax on imagination, and this particular tax invoice will be enormous. My position is simple: I will not underwrite any AI narrative that lacks a provenance audit, and I will not short the publishers. The trade that survives this cycle is not the one that feeds inside the legal gray zone. It is the one building the rails out of it. Impermanence is the only permanent yield, and in this trade the permanent asset is the demand for proof. The only question that remains is whether the frontier labs recognize that demand before the courts teach them the price.