Related works: Anthropic, Alibaba, and the Runtime Theft Problem | The Runtime Export Problem | Aerospace's Warning to AI, How Capability Laundering Will Reshape Corporate Compliance | The TSMC China License and the Limits of Hardware Export Controls
I. Executive Summary
Enforcement against AI extraction is about to feed the market it targets. Once regulators and deputized providers raise the price of direct querying, capability transfer migrates from runtime access to commerce in harvested behavioral corpora. The regulated object changes a third time: from the model, to the querying campaign, to the dataset the campaign produced.
Purchasing a corpus inherits the capability without the conduct. Anthropic’s September 10, 2026 threat report documented the market already operating. Proxy services log user exchanges with Claude and sell the transcripts, SenseTime purchased harvested corpora for training, and MiniMax built a shell proxy offering only American models. Two days earlier a tri-agency advisory from the NSA, CISA and FBI named six Chinese labs for the same conduct, so the extraction record moved from private allegation to convergent government attribution inside one week. The federal advisory and the drafted statute both stop one layer short of the trade.
Every instrument the operationalization turn produced assumes the conduct touches the provider’s runtime. Entity designation, intermediary categories and telemetry duties all watch the interface. A corpus transaction touches nothing a provider can see, so enforcement success at the account layer reads as progress even while capability transfer shifts to the one channel no telemetry observes.
The near-term equilibrium prices access, actors and account infrastructure while leaving the corpus itself unnamed. The artifact becomes strategically important before it becomes a recognized regulatory object. Provenance emerges first as private compliance and detection rather than as a named government category.
Extraction now produces a durable asset. A harvested corpus survives termination of the accounts, proxies and interfaces that generated it, so closing the access channel leaves the asset intact. Each transfer severs provenance further, until the downstream capability no longer traces to the conduct that produced it and the sequence runs from model to conduct to corpus to derivative capability.
Deputization supplies the displacement mechanism. Enforcement changes the composition of capability acquisition more reliably than it changes the total volume, so stronger account defenses can reduce extraction overall while shifting the surviving demand farther from the provider boundary. Silent downgrades, identity gating and reporting duties raise the price of direct extraction while leaving transcript purchase untouched, so the remedy for direct extraction raises the relative value of the substitute channel.
MindCast reads the contest through Predictive Behavioral Economics + Dynamic Game Theory, released through the MindCast AI Proprietary Cognitive Digital Twin Foresight Simulation (MP CDT FS) framework. Behavioral economics supplies the decision rule: institutions regulate the conduct they can observe, so salience anchors enforcement to the visible channel. Game theory supplies the payoff structure: rational demand routes to the lowest-penalty channel, and the predicted behavior emerges from the two together.
The June analysis in Anthropic, Alibaba, and the Runtime Theft Problem predicted enforcement would migrate from litigation to export control three months before three agencies and a provider converged on the September record. The Runtime Export Problem carries six live Simulation Predictions, and the events that settle them gate this paper’s release.
The paper maps the market’s supply chain and tests four competing claims to the corpus. The venue contest and the candidate rule follow, with the formal simulation register to come.
The formal MP CDT FS run releases eight Simulation Predictions, four Primary and four Secondary, each an event forecast with its band and with falsifiers and settlement sources in Section VII:
RD-P1 (74%, band 67–80): No named corpus or broker category enters any federal export instrument by September 30, 2027.
RD-P2 (78%, band 72–84): Private provenance and diligence practice develops before any government corpus category, by September 30, 2027.
RD-P3 (78%, band 70–87, conditional): The first broker enforcement action proceeds under non-export authority, by September 30, 2027.
RD-P4 (80%, band 73–87): A further authoritative publication documents intermediated or harvested-corpus acquisition, by September 30, 2027.
RD-S1 (80%, band 73–86): Buyer-side origin diligence appears before any seller-side broker licensing, by September 30, 2027.
RD-S2 (62%, band 53–70): A provider or government publicly discloses a source-detection event, by September 30, 2027.
RD-S3 (83%, band 76–89): Beijing does not burden its own labs’ use of foreign-harvested corpora, by June 30, 2027.
RD-S4 (72%, band 64–79): A federal measure prices the account or proxy channel, by September 30, 2027.
RD-S4 and RD-P1 jointly test the displacement asymmetry: the account channel acquires a price while the corpus channel stays unnamed. RD-P4 carries the migration observable, and RD-P2 carries the central finding, that private provenance practice organizes ahead of any government corpus category.
🏛️ Policymakers (RD-P1, RD-P2, RD-S4): The enforcement architecture now forming cannot observe the transaction that matters next. A corpus instrument designed before the migration completes beats a broker category named after the trade has scaled.
💼 Executives (RD-P4, RD-S2): Provider telemetry ends at the interface, and the corpus market operates past it. Detection investment that stops at account signals leaves the destination channel dark, so provenance marking belongs in the same budget line as extraction detection.
⚖️ Counsel (RD-P3, RD-S1): Clients now sit on three sides of one trade: providers whose outputs are harvested, enterprises whose sessions are logged for resale, and purchasers whose training data carries unknown provenance. Origin documentation separates licensed corpora from harvested ones before any rule requires the distinction.
📊 Investors (RD-P1, RD-P4): Capability moats priced on training cost were already mispriced against querying-based extraction. A corpus channel transfers the same capability at lower cost and lower legal exposure, so moat duration depends on the corpus channel alongside technical decay and model improvement, not on enforcement alone.
II. From Runtime Extraction to the Runtime Data Market
The corpus trade already has a supply chain, and the September record names each stage. Harvest, aggregation, sale and ingestion each appear in documented form. Distinct actors hold distinct exposures at every stage.
Loggers sit at the harvest stage. Anthropic documents proxy services that route user requests to Claude, log the full exchange, and retain the transcripts for resale. Users received their answers and never learned their sessions had become inventory.
Shells are loggers built for the purpose. MiniMax operated a shell company offering access only to Anthropic and OpenAI models, a storefront whose product was the traffic it observed. A shell runs little or no genuine service business, so the traffic it carries is available for harvest.
Brokers hold the aggregation and sale stage. Purchased transcript corpora appear in the record as a traded input, which requires a counterparty packaging harvested exchanges into training-ready datasets. Broker identity remains the least documented link in the chain, and the evidence program in Section VIII targets the gap.
Buyers close the chain at ingestion. SenseTime, tracked in the September report as campaign GTG-16012, purchased transcripts of user exchanges with Claude from third-party data vendors rather than generating fraudulent traffic itself. The purchaser inherited capability produced by conduct it never performed, which is the property that makes the channel strategically distinct.
Specialization drives the chain’s growth. An intermediary that solves the access problem amortizes that capability across many buyers, an aggregator raises corpus value through filtering and packaging, and a broker separates sellers from buyers who never meet. Each specialization improves market efficiency while lengthening the provenance chain by another link.
Digital replication defeats chokepoint enforcement. Seizing one shipment removes one copy and terminating one broker leaves the inventory reproducible, since a corpus can be split, recombined or transformed before resale. Interdiction strategies built for physical supply chains inherit an attribution problem and a replication problem at once.
The record is corroborated across venues rather than resting on one disclosure. OpenAI’s February 2026 memo to the House Select Committee described DeepSeek circumventing access restrictions through obfuscated third-party routers. By April the House record showed OpenAI, Google, Anthropic and xAI all reporting distillation pipelines against their models.
The trade runs in both directions, and the inbound flow complicates every clean narrative. Moonshot, DeepSeek and Xiaomi silently rerouted or replayed their own customers’ sessions through Claude, and the relayed traffic spanned at least a dozen languages. Relayed material included surveillance-footage analysis linked to the People’s Liberation Army and live credentials for a Russian defense-ministry database.
Beijing’s exposure therefore cuts both ways. Harvested corpora contain Chinese user sessions, and PRC data-security and personal-information law restricts cross-border movement of exactly that class of data. A PRC lab importing harvested corpora may face domestic legal friction Washington never engineered, a hidden variable no US-centric analysis models.
Four roles, one chain, and only one role named in any drafted instrument. The account-network provider carries statutory definition and designation machinery, while the logger, the broker and the buyer of harvested corpora operate outside every named category. One named role and three unnamed ones mark where enforcement can reach and where the market operates without legal friction.
III. Enforcement Displacement and the Emerging Supply Chain
Deputized enforcement converts frontier providers into monitoring nodes, and the conversion carries a structural blind spot. Advisory AA26-251A asks providers to flag accounts by behavioral signature, serve suspected accounts downgraded models, and share indicators across the industry. Every one of those duties executes at the runtime interface.
The transcript market operates past that interface. A logger harvests sessions inside lawful-looking traffic, a broker aggregates and sells the corpus, and a purchaser trains on it without ever touching the provider’s systems. The entire deputized apparatus watches a perimeter the decisive transaction never crosses.
Differential pricing follows directly. H.R. 8283 names the fraudulent account-network provider and attaches designation machinery to the role, while no instrument names the transcript broker. Account fraud carries designation risk, and corpus purchase carries none.
Dynamic game theory converts the asymmetry into a forecast of demand. Extraction demand faces two channels with one payoff and two penalty schedules, so demand shifts toward the unpriced channel as the priced one grows costly. Enforcement changes the composition of capability transfer rather than its volume, so total extraction can fall while the corpus channel gains share, bounded by depreciation of harvested traces against advancing models.
Behavioral economics explains why the blind spot persists. Account fraud is salient because providers see it, agencies can name it, and a statute already describes it. A corpus sale is invisible to every institution currently assigned to the problem, and institutions regulate what they can observe.
The advisory’s own remedy sharpens the effect. Serving suspected accounts a covertly downgraded model degrades the quality of directly extracted data. Degrading direct extraction raises the relative value of clean corpora harvested from unsuspected ordinary users.
Sunk inventory sharpens the displacement. Provider defenses operate prospectively while accumulated corpora represent completed extraction, so tightening access raises the scarcity value of every collection already outside the boundary. The effectiveness of future controls partly depends on how much transferable behavioral data escaped before the controls arrived.
June’s analysis argued attribution cost decides venue, and September’s argued operationalization decides instrument design. The displacement finding extends the sequence with a third proposition: instrument success at one layer changes the mix of demand at the next.
Provenance governance carries a second-order effect the run surfaced. Requiring source attestation improves visibility into the compliant supply chain while raising the strategic value of provenance laundering in the illicit one, so the same rule that organizes the legitimate market sharpens the incentive to obscure origin in the market it targets.
IV. The Property Problem: Who Controls a Behavioral Demonstration
A harvested corpus implicates at least four competing claims, and no legal regime was designed to allocate rights among them. The provider generated the outputs and the user authored the prompts, while the logger compiled the collection and the purchaser holds possession. Each claim draws on a different body of law, none written for the object.
The object resists every existing category. A behavioral demonstration is one exchange, and the strategic value of a corpus emerges statistically across millions of them. A single transcript maps awkwardly onto trade-secret, controlled-item or protected-work categories on its own, while the aggregate transfers capability that cost billions to create.
Copyright’s application is contested at every node of the object. Model outputs sit in contested authorship territory, user prompts are individually thin, and a compilation right would attach to the logger who assembled the harvested collection. Trade secret law fits the training signal better, yet the secret was extracted through millions of individually ordinary interactions that existing misappropriation theories describe awkwardly.
Contract reaches the account holder and stops there. Provider terms of service prohibit training on outputs and bind the account holder, while the broker and the downstream purchaser never signed anything. Contractual remedies reach the harvest but not the resale.
Coase’s framework organizes the problem. Where rights are undefined and transaction costs are high, possession allocates the resource, and the corpus market is a possession-allocation system operating at industrial scale. No bargaining among the four claimants can occur, because three of them do not know the transaction happened.
Becker’s framework prices the broker’s decision. Expected penalty is the product of detection probability and sanction, and with no named category, sanction is undefined and detection is rare. Expected penalty runs low and margin dominates, so broker growth is sustained but bounded, with Simulation Prediction RD-P4 carrying the observable.
Undefined rights are the condition that lets the market operate. Until an instrument defines who holds which right in a behavioral demonstration, the trade allocates capability to whoever harvests first and pays least. Defining the right is the precondition for controlling the transfer.
V. The Consent Problem: Users as Unwitting Feedstock
Every harvested corpus contains people who never consented to appear in it. Proxy users asked questions, received answers, and became training data for a foreign lab through a logging layer they could not see. Relayed sessions spanned at least a dozen languages, so the feedstock population is global while every named enforcement instrument is American.
Enterprise sessions carry two strategic values at once. A logged exchange reveals information about the user and demonstrates the behavior of the model, so one harvested corpus serves espionage and distillation from the same file. Proprietary code, internal analysis and research questions travel alongside the reasoning patterns a rival lab wants.
Consent doctrine gives regulators their fastest available hook. Logging sessions for resale without disclosure presents a conventional unfair-and-deceptive-practices theory under Federal Trade Commission (FTC) Section 5 authority, and state privacy statutes reach sale of personal information directly. Coverage turns on the facts, the representations and the specific statute. The theory is a route to a first action rather than a settled result. No new statute is required for the first enforcement action against a logger.
A venue race follows from the asymmetry in readiness. Export authority lacks a category, while consumer-protection authority holds one and needs only a target. Conditional on any action filing, the first enforcement action proceeds under non-export authority and enters the register as RD-P3, and whichever regime moves first defines the object for every regime that follows.
The race carries a consequence for Commerce. If the first precedent characterizes a harvested corpus as a privacy violation, BIS inherits an object defined by consent rather than by capability transfer. A privacy definition measures harm to the person in the transcript, and the national-security harm runs through the aggregate that no individual’s claim describes.
Privacy law therefore prosecutes the harvest and misses the transfer. A logger can settle an FTC action, adopt a disclosure banner, and keep selling corpora whose capability value is untouched by any consent remedy. Consent violations are the easiest feature of the market to prosecute and the least connected to the capability harm.
The consent problem still does real analytical work, because it identifies the enforcement entrant most likely to move first. FTC and state attorneys general belong in any model of this market alongside BIS and Congress, and their incentives point at the logger while the capability problem sits with the buyer. The first mover defines the object, and nothing requires the first mover to be thinking about export control at all.
VI. The Regulatory Object: Defining the Corpus Without Controlling Data
Export doctrine classifies items, software and technology, and a harvested corpus maps poorly onto each category. The difficulty is classification and administrability under existing definitions rather than categorical legal impossibility. The Export Administration Regulations (EAR) attach controls to defined objects and listed end users, while a corpus is data whose strategic significance is statistical and invisible at the single-record level, so controlling data as a class would sweep ordinary commerce while controlling nothing leaves the documented channel open.
Deemed-export doctrine shows the problem has been solved once before. Release of controlled technology to a foreign national is treated as an export to that person’s home country wherever the disclosure occurs, so export law already regulates an intangible transfer defined by content and recipient rather than by a border crossing. US Outsourcing, What Leaves America’s AI-Quantum Buildout When the Megawatts Stay maps the doctrine’s current operation, and the corpus problem is the same shape one object over: the megawatts stay while capability moves, the model stays home while capability leaves in data.
Five governance architectures now compete for the object, and each solves one problem while exposing another. Conduct governance targets the acquisition method and fails once the data leaves the extractor. Counterparty governance restricts designated actors and rewards intermediation.
Corpus governance defines regulated dataset classes and risks sweeping legitimate data commerce into a national-security architecture. Provenance governance requires origin attestation and stays vulnerable to transformation and commingling. Downstream-use governance reaches the training pipeline and demands visibility into practices governments do not possess.
The administrable answer is a hybrid rather than a winner among the five. Three definitional axes drawn from the competing architectures could bound a controlled corpus without controlling data.
Provenance separates harvested collections from licensed ones. Structure applies the H.R. 8283 totality indicia at dataset level, reading volume, capability concentration and development-timeline correlation as properties of a corpus rather than of a campaign. Recipient class attaches the control to covered entities rather than to the data itself.
Provenance is the binding informational constraint on administrability, and the technical contest is already live rather than prospective. Google DeepMind ships SynthID-Text watermarking in production, and published research on watermark radioactivity shows student models inherit detectable signatures from watermarked training traces, including reasoning traces. Robustness studies contest signature survival under paraphrase and mixing, so detection and laundering race, and the race recreates a familiar structure one layer down.
Zhipu abandoned an extraction attempt when Fable’s safeguards degraded the product and rerouted toward weaker interfaces, which established that deterrence in this domain is engineered rather than decreed. A detectable corpus is the artifact-layer version of the same finding. Provenance marking converts an unpriceable market into a priceable one, and no rule can administer what no party can detect.
A candidate rule assembles from the three axes together. Private provenance practice develops first, a counterparty duty is the more administrable public step, and a comprehensive corpus category is least reachable because it fails the classification test. The reachable public rule would control transfer of harvested-provenance corpora to covered entities while leaving licensed commerce untouched, and Simulation Prediction RD-P2 prices the private provenance layer developing ahead of any government corpus category.
The design borrows the deemed-export insight, the H.R. 8283 conduct test and the engineered-deterrence finding. Each component has an analogue in current law or practice, though the assembled regime does not yet exist.
Provenance severance is what the contest ultimately decides. Every aggregation, translation and paraphrase weakens the evidentiary link between the final dataset and its source, and after training the corpus disappears as an object while its value persists as derivative capability in another model. The governing question migrates from what item moved to what capability the recipient acquired, through what chain, and whether the chain can still be reconstructed.
What does not exist is the decision to assemble them, and S-3 predicts the assembly does not happen inside the next completed rule cycle. The gap is not ignorance, since the advisory demonstrates the government can describe the market. The gap is that every assembled component belongs to a different institution, and no venue currently owns the object.
VII. MindCast Foresight Simulation Predictions
Prediction ledger. Prediction date and data cutoff: September 12, 2026. The predictions below were fixed before the events they forecast. Each Simulation Prediction carries an event probability with a sensitivity band, a settlement window, an explicit falsifier and a named settlement source. Bands are event probabilities and are never aggregated.
Simulation synthesis. The dominant mechanism is differential channel pricing under telemetry blindness: enforcement prices the account channel while no instrument or sensor reaches the corpus channel. Enforcement changes the composition of capability acquisition more reliably than it changes the aggregate volume, so total extraction can fall while the corpus channel gains relative share.
The active regime is conduct and entity enforcement at the access layer. Near-term enforcement reaches access, actors and account infrastructure while the corpus stays legally unnamed. Provenance appears first as private compliance and detection, and government provenance or corpus categories do not necessarily follow inside the window.
Among the four candidate regulatory equilibria, the conduct and counterparty-dominant hybrid governs the window. Comprehensive corpus control fails the classification test that separates strategically meaningful harvested demonstrations from legitimate data. The corpus-sensitive equilibrium is not reached inside the window, and source-tracing governance emerges as private practice rather than as a named federal requirement.
VII.A Primary Simulation Predictions
RD-P1. No federal instrument names a transcript-broker or harvested-corpus category by September 30, 2027 (74%, band 67-80). No federal export-control rule or enacted statute in the window creates a named regulatory category for brokers of harvested inference transcripts or for harvested behavioral corpora as a controlled class. Falsified by any Federal Register rule or enacted statute that names and burdens the role or the dataset class. Settlement source: operative text in the Federal Register and enacted statutes. RD-P1 extends the inherited S-3 register, which prices the next single instrument at 79%; RD-P1 spans every instrument through the window, and the broader falsification surface prices it lower.
RD-P2. Private source-tracing practice develops before any government corpus category, by September 30, 2027 (78%, band 72-84). Private provenance and counterparty-diligence practice, provider output marking, buyer source attestations, dataset-origin representations or industry standards, materially develops before the United States creates any government category controlling harvested interaction corpora. Falsified if a government corpus category arrives first, or if no material private provenance practice develops through the window. Settlement source: published provider policies, industry standards, dataset-acquisition contracts and government rules. RD-P2 carries the central finding: the market develops source-tracing as private practice before government names the object.
RD-P3. The first transcript-market enforcement action invokes non-export authority, by September 30, 2027 (78%, band 70-87, conditional). If any federal or state enforcement action is filed against a transcript logger, broker or proxy reseller in the window, its lead authority is consumer-protection, privacy, fraud, computer-misuse or trade-secret law rather than export-control or sanctions authority. Falsified by an export-control or sanctions action arriving first. The register settles void if no qualifying action files in the window. Settlement source: the charging or complaint documents of the first filed action.
RD-P4. Further authoritative documentation shows intermediated or harvested-corpus acquisition, by September 30, 2027 (80%, band 73-87). At least one additional authoritative publication, a frontier-provider threat report, government advisory, indictment or congressional finding, documents capability-acquisition campaigns exhibiting intermediation, previously harvested outputs or third-party access that separates the ultimate beneficiary from direct provider interaction. Falsified if direct first-party extraction remains overwhelmingly dominant with no material increase in intermediated acquisition despite stronger provider defenses. Settlement source: the qualifying publications. RD-P4 carries the migration observable, subject to one limit: absence of new documentation could reflect either a real decline in the channel or the difficulty of observing it.
VII.B Secondary Simulation Predictions
RD-S1. Buyer-side origin diligence appears before any seller-side broker licensing, by September 30, 2027 (80%, band 73-86). Major legitimate AI actors exposed to corpus-origin risk adopt contractual provenance representations, audit rights, restricted-source clauses or equivalent buyer-side controls before any government licensing system for AI interaction brokers exists. Falsified if a seller-side broker licensing regime arrives first, or if no such buyer-side controls appear in the window. Settlement source: published contracts, provider policies and industry-standard documents.
RD-S2. A provider or government publicly discloses a source-detection event, by September 30, 2027 (62%, band 53-70). At least one frontier provider or government body publicly identifies watermarked, canaried or signature-bearing provider outputs inside a third party’s training data or model. Falsified if no such disclosure appears in the window. Settlement source: provider and government publications.
RD-S3. Beijing does not burden its own labs’ use of foreign-harvested corpora, by June 30, 2027 (83%, band 76-89). No PRC ministry or regulator in the window publicly burdens domestic labs’ acquisition or use of foreign-harvested interaction corpora under PRC data-security or personal-information authority. Falsified by any official measure or statement imposing such a burden. Settlement source: MOFCOM, CAC and MIIT publications.
RD-S4. A federal measure prices the account or proxy channel, by September 30, 2027 (72%, band 64-79). At least one federal measure, an Entity List addition, sanctions designation, indictment or named enforcement action, targets a fraudulent account-network provider, proxy operator or advisory-named distillation actor. Falsified if no such measure lands by the window close. Settlement source: BIS Entity List, OFAC Specially Designated Nationals list and DOJ dockets. RD-S4’s window runs past the September 24 summit, so it coexists with the inherited P-3 register on pre-summit timing.
RD-S4 and RD-P1 jointly test the displacement asymmetry at the register level: the account channel acquires a price while the corpus channel stays unnamed. RD-P4 carries the migration observable that the asymmetry predicts, and RD-P2 predicts the administrable form the response takes.
VII.C Risk Mitigation
Each registered Simulation Prediction below carries four parts. The claim and its band come first, then the exposure in the unit the stakeholder controls. Unilateral mitigating actions follow with an owner and a deadline, and the residual that survives full mitigation closes each entry.
Severity and probability are separate axes, so a lower-probability register with severe exposure can warrant more mitigation spend than a higher-probability register with contained exposure. Bands are event probabilities and never aggregate across registers. Actions are analytic options rather than legal, investment or fiduciary advice.
RD-P1. No named corpus or broker category. 74% (67-80). Severity: high. Binds Policymakers and Counsel.Exposure for policymakers: an enforcement gap sized by the volume of capability transfer moving through a channel no federal instrument names, growing each quarter the category stays absent. Exposure for counsel: a corpus purchaser holds datasets whose future controlled status is undefined, so acquisition terms written now carry unpriced retroactive risk. Actions: drafting staff scope the first reachable public duty as a counterparty obligation rather than a corpus definition, since the classification problem defeats the broader rule (owner: agency liaison, before the replacement-rule comment window closes). Counsel builds a data-origin record on every corpus acquisition now, ahead of any duty attaching (owner: compliance counsel, this quarter). Residual: an origin record still faces the attribution limit, since harvested corpora resist proof of source even when the buyer wants to establish it.
RD-P2. Private provenance practice precedes any government corpus category. 78% (72-84). Severity: medium. Binds Executives and Counsel. Exposure for executives: a provider that ships provenance marking late loses the window in which its marks become the de facto standard, measured in the share of downstream corpora its scheme can later identify. Exposure for counsel: contracts signed before provenance norms harden lack the representations that later become standard, forcing renegotiation at the next cycle rather than at signing. Actions: provider security and policy leads set a watermark and canary deployment date together with a disclosure policy, since selective release of a detection capability creates its own legal exposure. The constrained version is a pre-committed disclosure standard rather than case-by-case release (owner: security and policy leads jointly, this half). Transactional counsel adds source-attestation and audit-right clauses to the standard dataset-acquisition template now (owner: transactional counsel, before the next acquisition). Residual: private provenance improves the compliant supply chain while raising the strategic value of laundering in the illicit one, so the practice sharpens the incentive it targets.
RD-P3. First broker action proceeds under non-export authority. P(non-export lead authority given any action filed) = 78% (70-87). Severity: medium. Binds Counsel and Policymakers. Conditional register: the band applies only if an enforcement action files against a transcript logger, broker or reseller in the window; it makes no unconditional claim that an action files. Exposure for counsel: a logger or vendor client faces its live legal risk under consumer-protection and privacy law this year, not under the export regime its compliance program was built around. Exposure for policymakers: the first precedent defines the corpus as a privacy harm, and the capability-transfer characterization arrives at a venue that already owns the object. Actions: consumer-protection counsel treats FTC Section 5 and state privacy-statute exposure as the client’s operative risk and audits logging-disclosure practice against it (owner: consumer-protection counsel, this quarter); agency staff place the capability-transfer characterization on the record in any parallel proceeding before a first complaint fixes the framing (owner: interagency liaison, before the first qualifying action). Residual: venue selection belongs to whichever enforcer moves first, and no private party or single agency controls which authority reaches the object first.
RD-P4. Further documentation of intermediated or harvested-corpus acquisition. 80% (73-87). Severity: high. Binds Investors and Executives. Exposure for investors: a capability moat priced on training cost compresses as the intermediated channel persists, and the mispricing shows up as basis-point drift across every position underwritten on that moat. Exposure for executives: a provider whose threat model stops at first-party extraction under-measures its own capability leakage by the full volume of the intermediated channel. Actions: the research desk rebuilds moat-duration assumptions with an explicit channel-substitution term before the next underwriting cycle (owner: research desk, before the next commitment). The threat-intelligence lead instruments the next threat-report cycle to measure intermediated and purchased-corpus acquisition as a distinct category rather than a first-party total (owner: threat-intelligence lead, next reporting cycle). Residual: the channel’s true volume stays partly unobservable between disclosures, so any measurement is a floor rather than a full count.
RD-S1. Buyer-side origin diligence precedes seller-side licensing. 80% (73-86). Severity: medium. Binds Counsel and Executives. Exposure for counsel: a purchaser holding undocumented corpora inherits provenance liability the moment a duty or precedent attaches, and the discovery scope in any later dispute expands to every acquisition lacking an origin record. Exposure for executives: a training pipeline built on unattested corpora carries a capability whose legality cannot be established after the fact, which converts a completed model into a contingent liability. Actions: transactional counsel makes origin representations and audit rights a required term on dataset acquisitions now (owner: transactional counsel, before the next purchase); the data-governance lead conditions high-risk training-input purchases on source attestation at intake (owner: data-governance lead, this quarter). Residual: an attestation is only as strong as the upstream records behind it, so a diligent buyer can still inherit a laundered corpus that attests cleanly.
RD-S2. Public provenance-detection disclosure. 62% (53-70). Severity: high. Binds Executives and Counsel.Exposure for executives: a disclosed detection event converts corpus provenance from unpriceable to priceable overnight and reprices broker and dataset inventory a firm already holds. Exposure for counsel: a client holding corpora that a newly disclosed method can trace faces representations that were true at signing and false after disclosure. Actions: provider security and policy leads decide watermark deployment and the disclosure trigger together and pre-commit the disclosure standard, since selective disclosure creates its own exposure. The constrained version is a published standard rather than ad hoc release (owner: security and policy leads jointly, this half). Counsel for dataset purchasers adds a provenance-scan representation and a post-disclosure remedy clause to acquisition terms now (owner: transactional counsel, before the next acquisition). Residual: laundering research advances against every deployed scheme, so a detection capability priced today depreciates against tomorrow’s evasion.
RD-S3. No PRC burden on domestic labs’ use of foreign-harvested corpora. 83% (76-89). Severity: low. Binds Policymakers. Exposure for policymakers: the bidirectional narrative escalates in official rhetoric while the import channel stays legally open on the Chinese side, so a policy built on the assumption that Beijing will self-limit misreads the constraint. The cost is measured in analytic error carried into any measure that assumes PRC reciprocity. Actions: the China desk reads MOFCOM, CAC and MIIT output as two separate registers. Narrative and binding measure are scored apart, and any reciprocity assumption conditions on a binding measure rather than a statement (owner: China desk, before any measure premised on PRC self-limitation). Residual: an abrupt PRC enforcement turn against a domestic lab remains available to Beijing at any time and would falsify the register, so the low severity holds only while the channel stays open.
RD-S4. Account-channel actor designated, sanctioned or charged. 72% (64-79). Severity: high. Binds Executives and Investors. Exposure for executives: a designation or indictment reprices every counterparty relationship touching the named actor inside days, and the operational cost is the contract book that must be reviewed and repapered on a regulator’s clock. Exposure for investors: a position exposed to the six advisory-named labs or their intermediaries faces multiple compression on a sudden designation, since a naming event reprices supply and partnership relationships before the market re-rates them. Actions: the risk desk maps counterparty exposure to the six advisory-named labs and their intermediaries and pre-clears the repapering path before any designation lands (owner: risk desk, this quarter); portfolio risk stress-tests positions against a post-summit designation scenario and sets the position adjustment in advance (owner: portfolio risk, before the September 24 summit window). Residual: designation timing follows a political calendar no private model prices reliably, so the pre-cleared path reduces the scramble without removing the timing surprise.
Cross-register linkage. The origin-record actions under RD-P1 and RD-P2 also reduce RD-S1 and RD-S2 exposure, since one provenance record serves diligence, licensing readiness and detection response. The linkage is stated once and no action claims credit under more than one register.
VIII. Monitoring: Publication Triggers and Register Settlement
The paper is complete and holds for release until one of the forcing events below occurs. Each forcing event supplies the live newshook that opens publication, and each also settles at least one inherited Runtime Export Problem register. The inherited register released six Simulation Predictions, three Primary and three Secondary, restated here with frozen terms.
P-1 (84%, band 78-89). Through September 2027 federal policy operationalizes runtime extraction primarily through entity, end-user and intermediary measures and provider telemetry rather than licensing of ordinary foreign inference. Settlement source: BIS guidance, Entity List actions, State and OFAC measures.
P-2 (71%, band 64-78). The first Bureau of Industry and Security (BIS) rule replacing or materially implementing the rescinded diffusion framework contains at least one conduct-based element. Settlement source: Federal Register, RIN 0694-AJ90 or successor.
P-3 (83%, band 77-88). No Entity List addition or sanctions designation of the six advisory-named labs occurs before the 24 September 2026 Trump-Xi summit opens. Settlement source: BIS Entity List and OFAC SDN list.
S-1 (61%, band 52-69). Through the first quarter of 2027, Beijing escalates beyond neutral-technology framing and publicly invokes bidirectional exposure. Settlement source: MOFCOM, MOFA and MIIT statements.
S-2 (77%, band 69-82). Through September 2027, cross-provider extraction-indicator sharing formalizes into a publicly named standing mechanism. Settlement source: agency and provider announcements.
S-3 (79%, band 72-85). Through the next completed federal rule or enacted statute addressing model extraction, no named regulatory category is created for brokers of harvested inference transcripts. Settlement source: operative text of the first qualifying instrument.
S-3 is the register this paper’s thesis inhabits. The prediction expects the transcript layer to stay unnamed, and Sections IV through VI explain why naming it is genuinely hard rather than merely neglected. The analysis models non-naming as the expected route and treats naming as the falsifier branch, so the paper explains the gap without lobbying against its own register.
Five forcing events gate release. Each settles at least one register above and activates a section of this paper, and the first to occur becomes the publication newshook. The list doubles as the pre-release watch table.
Issuance of the diffusion-framework replacement rule. Settles P-2 on the conduct-versus-object question and materially updates P-1. Activates Section VI, since the rule’s text answers whether any corpus-layer element appears.
Committee movement or enactment of H.R. 8283. Advances S-3 settlement, since enacted text either names the broker role or confirms the omission. Activates Sections III and IV.
Entity List or sanctions designation of a named lab. Settles P-3 if pre-summit, updates P-1 regardless. Activates Section III, since designation raises the account channel’s price and starts the displacement clock.
Enforcement action against a proxy or reseller network. Resolves the venue race in Section V and supplies the first legal characterization of the object. Whichever statute grounds the action becomes the market’s first definition.
A second lab disclosure documenting transcript-market purchases. Corroborates the market’s scale and supplies the first possible measurement of channel migration. Activates Section II and the migration-scale prediction class.
Evidence closes resolved at the run of September 12, 2026. H.R. 8283 reported out of the House Foreign Affairs Committee 43-0 after an April 22 markup, with no broker category in the reported text. No federal enforcement action against a transcript logger, broker or reseller has been located to date.
OpenAI’s February 2026 House memo and subsequent industry-wide reporting corroborate the distillation record, while corpus purchase specifically remains single-source. SynthID-Text watermarking is confirmed in production, with watermark radioactivity validated in research. Still open: transcript-vendor identity and jurisdiction, any Senate companion, and jurisdiction-specific privacy analyses of relayed user data.
Signals map to the released register. The first qualifying instrument’s text settles the inherited S-3, updates RD-P1, and may settle RD-P2 if it carries a provenance duty. The first filed action against a transcript-market participant settles RD-P3.
The next threat-report cycle carries RD-P4 and RD-S2, while published contracts and provider policies carry RD-S1.
MOFCOM, CAC and MIIT publications carry RD-S3 through June 2027. The Entity List, SDN list and DOJ dockets carry RD-S4 through September 2027.
IX. Conclusion: The Artifact Is the Next Unit
Runtime Theft asked how many interactions transfer capability, and The Runtime Export Problem asked how telemetry becomes an administrable trigger. The corpus market answers both questions in a way neither paper’s instruments can reach: enough interactions to transfer capability now fit in a purchasable file, and no telemetry observes the purchase.
The unit of account has recursed twice. Commerce controlled the model and Congress drafted the campaign, and the corpus is now the strategic unit that holds the capability, still unnamed as a regulatory object. Enforcement reaches access, actors and account infrastructure while source-tracing develops first as private practice, so the state that recognizes the strategic unit early shapes the next enforcement cycle while the state that waits for a clean corpus category will still be drafting when the transfer has moved.
Deputization assigned providers the monitoring role, and provider visibility ends at the interface. Displacement now carries a priced register rather than an assumption: RD-S4 prices the account channel acquiring a penalty while RD-P1 prices the corpus channel staying unnamed. RD-P4 prices the documentation the shift leaves behind, and RD-P2 prices the private provenance practice that develops before any government corpus category.
The corpus is the last observable object in the chain. After training it dissolves into derivative capability, so the transaction is the last point at which the transfer can be seen. The behavior can end while the capability keeps moving.
The decisive signal is the first instrument from any venue that names the object. The statutory basis of that instrument will reveal whether the United States is regulating a privacy harm or a capability transfer, and the two definitions build two different futures.
Working With MindCast
MindCast converts institutional uncertainty into dated, falsifiable decision forecasts. Each engagement below keys to a register in Section VII.
Corpus-exposure audit. For counsel and compliance leaders at model purchasers and data vendors: an origin-documentation review separating licensed corpora from harvested-provenance risk, built against the dataset-level totality indicia before any instrument requires them.
Provenance-readiness assessment. For provider security and policy teams: a mapping of watermark and canary options against the laundering techniques the record documents, keyed to the provenance contest the formal run prices.
Corpus-instrument design brief. For policymakers and agency staff: the three-axis rule architecture in Section VI developed against a specific instrument, with the venue-race timeline and the privacy-characterization risk mapped before the first enforcement action defines the object.
Contact mcai@mindcast-ai.com. See Live-Fire Game Theory Simulators, Runtime Predictive Infrastructure.
Appendix A — Selected MindCast Works
The Runtime Arc
Anthropic, Alibaba, and the Runtime Theft Problem (2026). Established attribution cost as the venue-deciding variable and predicted migration to export control.
The Runtime Export Problem (2026). Documents attribution compression, names operationalization as the binding constraint, and carries registers P-1 through S-3 that this paper’s triggers settle.
Export-Control Lineage
Aerospace’s Warning to AI, How Capability Laundering Will Reshape Corporate Compliance (2025). Forecast compute-access licensing and access-layer governance before the reported rulemaking named that jurisdiction class.
The TSMC China License and the Limits of Hardware Export Controls (2026). Established that physical custody does not define the capability boundary, and ran the Chicago composite this paper reapplies.
China’s H200 Import Block and the Reordering of National Innovation Control (2026). Established the two-gate structure and the finding that neither gate governs post-delivery capability flow, the gap the corpus market exploits.
H200 China Policy Validation (2026). Documented the structural confirmations that validate the series architecture the runtime arc extends.
US Outsourcing, What Leaves America’s AI-Quantum Buildout When the Megawatts Stay (2026). Maps deemed-export doctrine’s regulation of intangible transfer by content and recipient, the closest existing analogue to a controlled corpus.
Strategic and Economic Overlay
The Global Innovation Trap (2025). Established capability capture at a fraction of originating R&D cost, the economics the transcript market perfects.
National Innovation Behavioral Economics (2025). Supplies the parent framework: institutions moving slower than the technologies they govern.
China Data Center Consolidation and H200 Exploit Pathway Evolution (2025). Predicted state-coordinated consolidation would raise exploit probability while lowering detection, the buyer-side structure of the corpus trade.
Why AI Commoditizes Raw Prediction, Why Governance Stays Scarce (2026). Supplies the governance-scarcity engine explaining why the extractor internalizes gains while boundary costs socialize.
The AI Duel of America’s Chaotic Advantage vs. China’s Disciplined Coordination (2025). Provides the comparative behavioral architecture behind the two-board framing.
The Beijing Summit Validation (2026). Establishes the two-board structure governing all China-side registers in the arc.
Conditional entry, per the planning brief: Pentagon-Anthropic Throughput Failure and the Structural Reclassification of Safety as Ideology (2026) enters only if the formal run finds government coercion of provider governance causal to the displacement mechanism.
Appendix B — External Sources
Primary record
“Detecting and Countering Misuse of AI: September 2026,” Anthropic, September 10, 2026.
“China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies,” Advisory AA26-251A, NSA, CISA and FBI, September 8, 2026.
H.R. 8283, Deterring American AI Model Theft Act of 2026, 119th Congress, introduced April 15, 2026.
Export Administration Regulations, 15 C.F.R. § 734.13 (deemed export provisions).
H.R. 8283 legislative record, Congress.gov; reported out of the House Foreign Affairs Committee 43-0 following the April 22, 2026 markup.
OpenAI, “Updated Stakes for American-Led, Democratic AI,” memo to the House Select Committee on Strategic Competition, February 12, 2026 (as reported by Reuters).
Secondary record
“U.S. AI Export Controls in 2026: A Practitioner’s Guide,” One Lex Partners, June 19, 2026.
“New US Export Controls Reportedly Target Chinese Access to Remote AI Servers,” Tom’s Hardware, July 22, 2026.
“AI Distillation Attacks: Executive and Congressional Action Can Go Further,” Institute for AI Policy and Strategy, May 2026.
Dathathri et al., “Scalable watermarking for identifying large language model outputs” (SynthID-Text), 2024; Sander et al. on watermark radioactivity, 2024; “TextSeal: A Localized LLM Watermark for Provenance and Distillation Protection,”2026; “Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?” 2025.



