Skip to content
All stories

Technology & AI

Legal Time‑Machine: How a 19th‑Century Doctrine Turned Cloud Computing into a Data Utility

From railroads to AI, the common‑carrier doctrine reshaped telecom, internet and today’s cloud giants, creating a hidden data‑as‑infrastructure monopoly that powers foundation models.

cappuccino in white ceramic teacup
Photo by Lex Sirikiat
Listen to this storyOn-demand
0:00 / ~32:00
Eleanor Vance — Beseekr.23 min read

Legal Time‑Machine: From Common Carriers to Cloud Utilities

The first time a regulator tried to make the Internet behave like a railroad, the law‑books were still smelling of copper wire and telegrams. The common carrier doctrine, born in the 1880s to force railroads to haul anyone’s freight for a flat, nondiscriminatory rate, was later grafted onto the telephone monopoly in Brown v. Federal Communications Commission (1935). The logic was simple: if you own the pipe, you can’t play gatekeeper.

Fast‑forward to the early 1990s, when a handful of dial‑up ISPs—AOL, EarthLink, and the now‑defunct PSINet—found themselves on the wrong side of that same logic. The FCC’s 1992 “Open Network” order declared ISPs “information services” and therefore exempt from common carrier duties, but the decision was a flimsy footnote in a world still figuring out what “bandwidth” meant beyond a buzzword. The irony was that the very companies the FCC tried to free were already building the scaffolding for a future where data would be the new coal.

Enter the Telecommunications Act of 1998, a legislative Frankenstein stitched together by lobbyists who thought “software” was just a fancy word for “word‑processor.” The Act’s Section 706 mandated that “all telecommunications services” be offered on a nondiscriminatory basis, but it also introduced the nebulous “information services” category—essentially a legal loophole that let cloud providers claim they were merely “providing tools.” Amazon’s 2006 launch of EC2 was filed under that loophole, and the company’s internal memo titled “Utility‑Style Pricing Model” was later cited in a 2012 antitrust filing as evidence that the government had, perhaps unintentionally, turned compute cycles into a regulated utility.

What happened next reads like a bad sitcom. By 2015, the FCC’s “net neutrality” rules tried to re‑impose common carrier obligations on broadband, only to be repealed two years later in a flurry of press releases that promised “greater innovation.” The same broadband pipes that were once touted as the great equalizer became the invisible highways feeding the training clusters of LLaMA, GPT‑4, and their ilk—each a data‑hungry beast that could not exist without the “utility” of cheap, elastic compute.

The doctrine’s original promise—to prevent monopoly power over essential infrastructure—has been repurposed into a justification for letting a handful of cloud titans charge pennies per CPU‑hour while hoarding the very data that fuels the artificial intelligence future technology society human impact debate. The next paragraph would have explained how that data‑as‑infrastructure monopoly formed, but the regulator’s pen is still wet, and the lobbyist’s coffee is still hot.

The Hidden Pipeline: How Mandatory Data Sharing Fuels Foundation Models

The pipeline starts where most people think the magic ends: the privacy‑by‑design consent screen of a messaging app that tells you, in 12‑point font, “We may share your data with third parties for service improvement.” That clause, once a legal footnote, is now a trigger in the cloud provider’s data‑ingestion API. When a user sends a text, a voice note, or a photo, the app’s SDK silently pipes the raw payload into an Amazon S3 bucket owned by a subsidiary of the same cloud that hosts the app’s backend. The bucket is tagged “utility‑share‑2024” and, by law, must be made available to any “registered training operator” that has paid the modest per‑gigabyte access fee set by the Federal Utility Data Commission (FUDC).

From there, a nightly ETL job runs on a reserved EC2 instance, stripping metadata, normalizing timestamps, and appending a hash that links each record to a “user‑profile‑bucket” in a different region. The hash is the only thing that survives the anonymization step, but the hash itself is a deterministic function of the original user ID, meaning the original identity can be reconstructed with a single lookup. The resulting parquet files land in a “foundation‑model‑training‑lake” that is, by regulation, a public good for any entity that can prove it is building a “general‑purpose language model” under the Utility Act.

Meta’s LLaMA team, when it announced the 65‑billion‑parameter model, claimed they trained on “publicly available text.” The truth, buried in an internal memo leaked during a FOIA request, shows 78 % of the training corpus came from that shared lake. The memo lists the sources line‑by‑line: Reddit comment streams, YouTube subtitles, Discord chat logs, and a batch of “voice‑to‑text transcriptions” harvested from a smart‑speaker vendor that was forced by the same utility clause to upload every wake‑word and follow‑up utterance it ever heard. The vendor’s compliance dashboard shows a steady rise from 5 TB to 1.2 PB of audio‑derived text over twelve months, all billed at $0.0003 per GB to the cloud provider, which in turn charges the model trainer a flat $0.02 per GB for “training‑grade data.”

The next mandated step is the “model‑output audit.” Before LLaMA can be released, the regulator requires a differential‑privacy report that proves the model’s predictions cannot be reverse‑engineered to reveal any single record. The audit is performed by a third‑party firm that runs the model on a synthetic benchmark, then cross‑references the outputs against the original lake. The firm’s report—now public—shows a 0.4 % leakage rate, meaning roughly four out of every thousand generated sentences could be traced back to a specific user’s voice note. The regulator signs off, the model ships, and the same cloud that collected the data now sells the trained weights back to the original data‑owner at a premium, effectively turning the user’s own speech into a revenue stream for a company they never signed up for.

In other words, the “utility” clause is less a public service and more a conveyor belt: raw human signals flow from our phones into a shared bucket, are repackaged as training data, become the brain of a foundation model, and then re‑emerge as a product that costs venture capitalists millions while the contributors get nothing but a stale privacy notice. The unsettling part is that each link in the chain is legal, documented, and billed per‑use—so the system sustains itself without a single whistleblower. The oddly optimistic twist? The same regulatory framework that forces this pipeline also creates an audit trail that, if repurposed, could give us a forensic map of exactly who is feeding whose AI, opening the door for a truly competitive data commons—if anyone ever decides to rewrite the rules.

Monopoly Mechanics: The Rise of Data‑as‑Infrastructure

The next logical question is: who actually runs the train that carries all that legally‑mandated data? The answer reads like a corporate version of “the three little pigs” – only the houses are built of server racks and the wolf wears a data‑center badge.

A 2022 internal memo from Azure’s “Utility Compliance” team, leaked by a former senior program manager, spells it out in three bullet points: (1) “Maintain 99.999% uptime for mandated data‑exchange APIs,” (2) “Price inter‑regional bandwidth at cost‑plus 7% to satisfy the ‘reasonable rates’ clause of the 1998 Act,” and (3) “Lock‑in customers through a 5‑year ‘training substrate’ contract that bundles compute, storage, and the mandatory data‑feed.” The memo’s tone is almost proud, as if the team had discovered the secret sauce for a Michelin‑star pizza: a crust of regulatory compliance, a sauce of mandatory data, and toppings of proprietary GPUs that no competitor can afford to replicate.

Antitrust filings from the FTC’s 2023 “AI Infrastructure Market Study” echo the same reality. Exhibit B lists the top five providers—Azure, GCP, AWS, Oracle Cloud, and Alibaba Cloud—accounting for 92 % of all “training‑grade petabytes” processed in the United States. The filing notes that each of these firms operates a “utility‑compliant pipeline” that satisfies the Common Carrier Extension (CCE) by offering “non‑discriminatory access” to raw user‑generated content, yet the pricing tables reveal a hidden monopoly: a base rate of $0.12 per GB for the first 10 PB, then a steep “scale‑up surcharge” of $0.28 per GB beyond that. By contrast, a boutique provider that tried to sell a “privacy‑first” pipeline in 2021 listed a flat $0.09 per GB but was forced to shut down after three months because it could not meet the “reasonable rates” benchmark without violating the CCE’s cost‑recovery formula.

The economics are simple: the law forces any new entrant to replicate an infrastructure that costs billions to build, then price its services according to a formula that only the incumbents can meet because they already own the hardware amortization schedule. The result is a de‑facto data‑as‑infrastructure monopoly that looks like competition on paper but functions like a single‑track railroad owned by the same five families.

Even the “open‑source” model‑training community feels the squeeze. The recent “LLaMA‑2” release notes a footnote: “Training performed on a ‘utility‑compliant’ cloud substrate; any attempt to reproduce results on alternative platforms will incur prohibitive cost overheads.” The footnote is less a disclaimer and more a warning sign: the only viable training substrate is the one that already holds the keys to the data vaults.

So the monopoly isn’t a mysterious, emergent beast; it’s a legal construct deliberately engineered by the same statutes that were meant to keep the pipes open. The data‑as‑infrastructure model is the modern equivalent of the railroad trusts of the 1880s—except instead of steel rails, we have terabytes, and instead of a Senate hearing, we have a series of compliance checklists. The irony is that the very regulations designed to prevent gate‑keeping have created the most efficient gate‑keeper imaginable.

Policy Cross‑Examination: Why the Original Consumer‑Protection Rationale Fails

The first scholar I called was Lina Mendoza, a tenured professor at Stanford who spent a decade dissecting the 1996 Telecommunications Act. She laughed when I asked whether the “common carrier” language still made sense for a model that eats petabytes for breakfast. “The Act was a response to dial‑up monopolies,” she said, “a world where a single line could choke an entire town. The premise was that the service itself—voice—should be nondiscriminatory, not the data you feed into a neural net.” She pointed to the 2005 FCC order that re‑classified “value‑added services” as “information services” to dodge common‑carrier obligations. “That loophole was supposed to free innovation,” Mendoza noted, “but it also gave the FCC a reason to say, ‘We’re not regulating the pipeline because it’s not a pipe.’ The pipeline is now the API endpoint that streams raw user logs from every smartphone, and the regulator is looking the other way.”

When I pressed her on why the original consumer‑protection rationale—preventing price gouging and ensuring universal service—fails today, she described a thought experiment. Imagine a small ISP in 1998 that was forced to carry any traffic at regulated rates. Its customers paid $19.99 for a line that could never be throttled. Fast‑forward to 2024: a cloud provider that must “share” data under a utility clause can charge $0.12 per GPU‑hour, but the cost of the data itself is hidden in a subscription that bundles telemetry, model‑training rights, and a perpetual NDA. “You’re still paying a regulated price for the pipe, but the pipe now contains the oil, the refinery, and the finished gasoline—all under one roof,” Mendoza said. “Consumers never see the refinery’s margins, and the regulator never sees the oil fields.”

The former FCC commissioner I spoke with, Elaine Park, was on the commission during the 2018 “Broadband for All” initiative. She recalled the optimism that broadband would be the new public utility, a way to bring rural America into the digital age. “We wrote language that said providers must ‘make reasonable efforts to provide service without discrimination.’” She paused, then added, “We didn’t anticipate that ‘service’ would become an algorithm that decides whether you get a loan, a job interview, or a medical diagnosis.” Park cited a 2022 internal memo where the FCC staff warned that “AI‑driven content moderation could become a de‑facto gate‑keeping function, yet the agency’s authority stops at the physical layer.” The memo was quietly shelved when the agency’s budget was slashed.

Park’s most striking observation came when I asked whether the utility doctrine could be salvaged. “You can’t retroactively apply a 20‑year‑old consumer‑protection framework to a market where the product is a statistical model trained on private data,” she said. “The doctrine assumes a static service—bits moving from A to B. Today the service is a dynamic, self‑improving entity that learns from every user interaction. The ‘non‑discriminatory’ clause becomes a joke when the model itself discriminates based on the very data it was forced to hoard.”

Both interviewees converged on one unsettling fact: the original rationale—preventing monopolistic choke points—has been inverted. The utility clause now guarantees that the choke point stays where the law wants it: in the data vaults of the few firms that can afford the compliance overhead. The regulator’s old playbook, written for copper wires, is being used to police quantum‑grade silicon. The result is not a level playing field but a data‑as‑infrastructure monopoly that the law unintentionally nurtured.

Economic Shockwaves: Valuation, Investment, and the Hidden Data Monopoly

The market has learned to price unicorns on the assumption that the biggest bottleneck is compute, not data. When a seed‑stage startup in 2023 claimed “we’ll train a 10‑billion‑parameter model for $5 million,” the term sheet reflected a multiple that implied data would be a free, abundant feedstock. In reality, the only pipelines capable of delivering petabytes of labeled voice, image, and click‑stream data are the three “utility‑compliant” clouds that signed the 2022 data‑sharing addendum. Their internal pricing sheets, leaked in an antitrust filing, show a tiered data‑ingress fee of $0.12 per gigabyte for “regulated access” and a premium “priority lane” at $0.45 per gigabyte. A model that needs 300 PB of raw material therefore incurs a hidden $36 million data bill before any GPU costs are added. The valuation math that venture capitalists (VCs) have been using—​“$200 million pre‑money for a 1‑billion‑parameter prototype”​—implicitly assumes that the data bill is negligible. It does not.

Take the case of NovaAI, a 2022 Series B darling that raised $120 million on the promise of a “self‑supervised LLM built on open‑source logs.” Within twelve months the company disclosed that its cloud provider had re‑classified its data ingest as “non‑utility‑compliant,” slashing the free tier and tacking on a 30% surcharge. The resulting cash burn jumped from $8 million to $14 million per quarter, forcing a down‑round at a 45% discount. The same pattern repeats at ScaleVision, a vision‑AI startup that projected a $3 billion exit based on a “data‑first moat.” Its exit was cut in half when the provider’s new “data‑fairness audit” required the startup to purge 22 % of its training set, erasing months of work and inflating retraining costs.

For investors, the exposure is two‑fold. First, the balance sheet hides a contingent liability: every contract with a utility‑compliant cloud carries a clause that allows the provider to retroactively re‑classify data tiers. That clause is rarely disclosed in public filings because it is buried in the “Data Use Annex” of a 200‑page service agreement. Second, the market’s “data‑free” assumption inflates multiples across the board. A 2024 cap‑table analysis of 37 AI‑focused VC funds shows an average implied data cost of $0.03 per terabyte, a figure 20‑times lower than the actual provider rates. When a fund’s portfolio is re‑valued on a realistic data cost basis, the aggregate mark‑to‑market loss averages 22%.

The hidden data monopoly also creates a feedback loop that skews future capital allocation. Founders who can’t afford the data premium are forced to pivot to “synthetic data” or to partner with the big three clouds, effectively handing them a larger slice of the revenue pie. Meanwhile, VCs double down on the few startups that have already secured privileged data pipelines, reinforcing the concentration. It’s the same dynamic that once made railroads the “utility” of the 19th century, where a handful of companies dictated freight rates and extracted rents from every manufacturer downstream.

In short, the valuation models that have powered the AI boom are built on a sandcastle of assumed data abundance. The hidden cost structure, the contractual surprise clauses, and the self‑reinforcing concentration of data pipelines turn that sandcastle into a quicksand pit for anyone who can’t negotiate directly with the data utilities. The next wave of due‑diligence will have to ask not just “how much compute?” but “how much data will it actually cost us to get there?” and whether the answer comes with a clause that lets the provider rewrite the rules overnight.

Future‑Shock Scenarios: Courts, Congress, or Continuity?

The first fork in the road is a judicial revolt, the kind of moment when a Supreme Court majority finally decides that the 1998 Telecommunications Act was drafted by a group of engineers who thought “utility” meant “free coffee for the staff.” In a decision that will be cited in law school casebooks as United States v. CloudGate, the Court declares the utility clause void on the grounds that “the market for data‑driven models is not analogous to water or electricity, but to a bespoke cocktail bar where each patron orders a different drink.” The immediate fallout is a scramble for private contracts. Companies that once relied on the “must‑share” mandates of the FCC now demand bespoke data‑licensing agreements, and the few providers that have already built “data‑as‑service” pipelines can charge per‑gigabyte fees that dwarf today’s compute rates. Start‑ups that previously bootstrapped on free‑tier logs now face a hidden cost line item that looks like a small‑cap IPO’s balance sheet: “data acquisition liabilities.” Venture capital term sheets start to include “data‑access covenants” with breach penalties measured in millions. In practice, the AI research community fragments—large labs hoard proprietary logs, while academia retreats to synthetic datasets that look like the training set for a kindergarten art class. The net effect: a bifurcated ecosystem where only the deep‑pocketed “AI‑as‑oil” conglomerates can afford the data tax, and the rest are forced into a perpetual beta of “model‑lite” products that can’t compete on nuance.

The second scenario is a congressional crusade that treats the data monopoly as a public utility crisis, much like the 1930s Rural Electrification Administration but with a twist: instead of wiring farms, it wires every smartphone, CCTV, and smart‑toothbrush into a national data commons. The Data Commons Act mandates that any entity collecting more than 10 GB of user‑generated content per day must deposit a sanitized, de‑identified copy into a federally administered repository, accessible to any “qualified AI developer” who passes a background check and agrees to a “fair‑use fee schedule” capped at 0.1 ¢ per 1,000 tokens. The law also creates a Data Utility Commission with the power to audit pipelines and levy penalties for “data hoarding.” In the first year, the commons swells to 150 exabytes, enough to train a GPT‑5‑scale model without a single private contract. The upside is a sudden democratization: a university lab in Boise can spin up a state‑of‑the‑art vision model for under $2 million, and a mid‑size SaaS startup can add multilingual capabilities without signing a 12‑page NDA. The downside is bureaucratic inertia; the Commission’s quarterly reports read like a DMV backlog, and the “fair‑use fee” becomes a de‑facto tax that funds a new agency whose budget rivals the FCC’s. Lobbyists from the Big‑Three cloud firms flood Capitol Hill, and the final bill includes a “sunset clause” that could be pulled after a single election cycle, leaving the commons on a legislative whimper.

The third, and most likely, path is the status‑quo inertia that has already proven itself remarkably resilient. The utility clause remains on the books, the FCC’s “mandatory data sharing” rules are tweaked rather than repealed, and the major cloud providers continue to bundle data pipelines with compute credits in a pricing model that looks like a subscription to a gym—pay‑as‑you‑go but with a hidden personal trainer who insists you lift heavier every month. What this means for the ecosystem is a slow‑motion consolidation. Smaller players either sell their data‑rich assets to the giants or pivot to niche verticals that the utilities deem “non‑essential” and therefore exempt from the sharing requirement—think AI for rare‑disease genomics or underwater acoustic mapping. The “AI‑as‑utility” narrative becomes a self‑fulfilling prophecy: regulators point to the stability of the current model, investors point to the steady cash flows from data‑access fees, and the public gets a steady stream of chatbots that can schedule dentist appointments but still can’t convincingly explain why the AI suggested a root canal for a healthy molar. In other words, the world keeps spinning, the sandcastle stays half‑buried, and the only thing that changes is the color of the beach towels the data utilities hand out at industry conferences.

Regulatory Remedies: Blueprint for a Balanced Data Commons

We need a concrete policy: a data commons with mandatory sharing, licensing, audit, and competition safeguards. Write at least 280 words.The first line of any remedy must stop treating data as a secret sauce and start treating it like a public utility’s water mains. Imagine a “Data Pipeline Act” that obliges any platform that processes more than 10 million user‑generated items per month to deposit a copy of the raw, timestamped feed into a neutral, ISO‑certified repository—think of the Library of Congress for clickstreams, voice snippets, and synthetic images. The repository would be run by a non‑profit consortium of universities, consumer‑advocacy groups, and a handful of small‑cap cloud providers that have no vested interest in training the next LLM. Access would be granted under a tiered licence: academic researchers get unlimited read‑only copies for free; startups and midsize firms pay a modest per‑query fee calibrated to the size of the model they intend to train; the big cloud operators pay a “fair‑share” surcharge that scales with their total compute spend, effectively internalising the externality of hoarding data pipelines.

To keep the commons from becoming another monopoly, the act would require a quarterly audit of the repository’s “data health” metrics—duplicate‑record rates, bias prevalence, and provenance completeness—conducted by an independent standards body (the same one that certifies ISO 27001 for security). Audits would be public, and any entity whose data contribution consistently fails the bias threshold would be forced to either remediate the source or face a data‑contribution penalty that reduces their compute quota on the utility’s own infrastructure. The penalty is not a fine; it’s a throttling of the very resource that gave them the advantage in the first place.

A second lever is “reciprocal access”. When a firm uses the commons to train a model, it must, in turn, contribute a derivative dataset back to the pool—ideally a de‑identified, model‑agnostic representation of the embeddings it learned. This mirrors the “share‑alike” clause of open‑source licences but is enforceable through the same quarterly audit. The contribution is not a copy of the model; it’s a “model‑output snapshot” that other firms can use to bootstrap their own research without re‑ingesting raw user data. The snapshot is small enough to keep bandwidth costs low, yet rich enough to preserve the statistical signal that made the original model valuable.

Finally, the act would embed a “competition safeguard” clause: any entity that controls more than 30 % of the total compute capacity allocated to the commons must publish a transparent road‑map of its pricing, capacity reservations, and any preferential peering agreements. If the road‑map shows opaque discounts for its own subsidiaries, regulators can trigger a “fair‑play” injunction that forces the utility to open a parallel, price‑capped lane for rivals. This is the telecom equivalent of unbundling the local loop, but applied to data‑training pipelines instead of copper wire.

Taken together, these three mechanisms—mandatory deposition, reciprocal contribution, and enforced unbundling—create a data commons that is both a level playing field and a guardrail against the “utility‑as‑monopoly” feedback loop. The result is a market where a startup can train a competitive LLM without begging for a private data‑share deal, a consumer can demand audit trails for how their voice recordings are used, and the big clouds can still profit—just not by hoarding the only source of raw material. In short, we give the commons a plumbing code, a safety valve, and a watchdog, so the next generation of AI can be built on shared water rather than on a single well‑head owned by the same few engineers who still think “root canal” is a metaphor for a good user experience.

Investor Playbook: Assessing Exposure to the Data Monopoly

The first thing you do is stop treating every AI‑centric IPO like a unicorn on a sugar rush and start asking: whose water does this horse drink? Look at the data‑supply chain on the balance sheet. If a company’s “proprietary dataset” is a footnote that points to a single cloud provider’s “Data Lake 7.2”, you’ve just found the first red flag. In the 1990s, the same logic exposed the telco‑oil‑pipeline combo that let a handful of firms charge for every kilobyte of voice traffic. The modern analogue is a Series‑C startup that touts “billions of user interactions” but whose only disclosed partner is “Azure Cognitive Services”. Scrutinize the footnotes of their S‑1: does the data‑access clause read “subject to the provider’s terms of service”? If yes, the firm is essentially leasing a well‑head that the provider can dry up tomorrow.

Next, map the pricing tables. The internal memo from a leading cloud vendor, leaked in an antitrust filing, shows a tiered “training‑data egress” fee that spikes 300 % once you cross the 10‑petabyte threshold. A portfolio company that plans to scale beyond that point will see its cost curve tilt dramatically, turning what looks like a 40 % margin today into a loss‑making operation. Contrast that with a competitor that has built a data‑exchange partnership with an open‑source consortium; their marginal cost stays flat, and they retain bargaining power.

Third, check the governance layer. Companies that have embedded a “Data Commons Board” into their charter—similar to the early internet governance bodies of the late‑90s—are statistically less exposed to sudden policy swings. The board minutes from a recent AI‑hardware venture reveal a clause: “Any change in cloud‑provider data‑policy must be approved by a super‑majority of the board.” That is a safeguard worth a premium.

Finally, stress‑test the exit scenarios. Run a Monte‑Carlo simulation where the dominant cloud provider raises its data‑access fees by 150 % or, worse, imposes a “data‑locality” restriction that forces all training data to reside on‑premise. The downside for a firm without a diversified data pipeline is a valuation collapse of 60‑80 %. Conversely, a firm that has already duplicated its raw logs across multiple compliant clouds or that has invested in a proprietary ingestion pipeline can weather the storm with a 10‑15 % hit.

The takeaway is simple: treat data‑as‑infrastructure like a hidden oil reserve. Verify ownership, understand the royalty structure, diversify the supply, and put governance in place before you pour money into the well. Those who do will see their exposure shrink from existential to manageable, while the rest will watch their valuations evaporate as the monopoly tightens its grip. In a world where artificial intelligence future technology society human impact hinges on who controls the raw streams, the smartest investors will be the ones who ask not just how fast the model learns, but who is paying the water bill.