ALAN opinion 12 min read

The Ethics of LLM Cost Cutting: Quality Erosion, Access Inequality, and Who Bears the Hidden Costs

Tiered model quality scales depicting unequal AI access — premium capabilities on one side, constrained outputs on the other

The Hard Truth

A documented study found measurable quality gaps in AI-assisted college admissions essays tied directly to socioeconomic status. Not predicted. Measured. The routing systems that produced these gaps were designed by engineers optimizing for cost efficiency. The optimization worked exactly as intended.

The engineers who built this infrastructure are not villains. They responded to genuine pressure — inference costs that climb steeply enough to make AI services financially difficult to sustain at scale. The pricing tiers reflect real trade-offs. The algorithms route traffic with mathematical precision. And somewhere in that chain of defensible, well-intentioned decisions, the question of who gets routed to which model was answered quietly, without debate, by market forces that were never asked to care about equity.

That is the problem this essay is trying to name.

The Evidence That Didn’t Make the Press Release

Research published in the GenAI Divide paper documented measurable output quality gaps in AI-assisted college admissions essays — gaps correlated directly with socioeconomic status. Students with the resources to access premium models, to iterate across multiple sessions, to prompt more effectively — they produce better AI-assisted work. Students who lack those resources produce work constrained by the models available to them.

This is not speculation about AI’s future. It is a measurement of what the present is already doing. And it matters because college admissions is one of those domains where unequal inputs compound: a better essay does not simply help one student, it displaces another who wrote an equally genuine but less polished application. The assistance gap becomes a displacement mechanism.

A framework for understanding AI access identifies three distinct levels: access to AI at all; the ability to use it effectively; and whether that use translates into real outcomes — income, civic participation, professional opportunity. Cost optimization debates tend to happen entirely at the first level. The inequalities that compound do their work at the second and third, where they are far harder to see and far easier to blame on the individual rather than the infrastructure.

The Architecture of the Gap

LLM Cost Management is not a single tool. It is a constellation of practices — Model Routing, Token Budget controls, Prompt Caching, Semantic Caching — that together determine which requests reach which models. The practices are effective. RouteLLM, developed by researchers at UC Berkeley, Anyscale, and Canva, demonstrated that routing frameworks can reduce costs substantially while preserving the majority of premium model quality across the routing system as a whole (RouteLLM paper, ICLR 2025).

The problem is not that these tools exist. The policy is encoded in the objective function. When Model Tiering intersects with unequal user resources, routing allocates premium inference budget by proxies — usage volume, subscription tier, session length — that correlate with wealth without naming it. A student on a free tier, a nonprofit running on the cheapest available API, a researcher in a jurisdiction where credit card billing to Silicon Valley simply doesn’t work — all of them receive economy-class responses from infrastructure that presents itself as democratizing access.

Research on Bias Amplification adds a dimension the efficiency argument does not account for. Studies have found that ML model predictions amplify biases present in training data, generating outcomes that exceed what training statistics alone would predict (Hall et al., arXiv 2201.11706). If lower-resource users are systematically routed to models with weaker alignment work, they do not merely receive lower-quality outputs — they receive outputs more susceptible to the biases that higher-tier models have been more extensively trained to mitigate. The routing decision, ostensibly about cost, becomes a decision about whose reality the AI reflects most accurately.

The Accounting We Don’t Publish

There is another cost the efficiency argument tends to omit. Morrison et al. (2025) found that development costs — failed training runs, ablations, iterative refinement — account for over four-fifths of total compute in model development, a figure almost entirely unreported by model developers. The energy consumption of generative AI runs at roughly 29.3 TWh per year, comparable to a small nation’s total electricity use, with sixty to seventy percent of that drawn from serving predictions rather than training. These costs are distributed globally, landing on communities that had no voice in the construction of this infrastructure and receive none of its economic returns.

When we count the savings from tools like LiteLLM-based routing frameworks or caching pipelines, we are calculating savings against an API billing statement. The environmental and resource costs that never appeared in that bill remain invisible, exported to places the optimization spreadsheet cannot see.

Security & compatibility notes:

  • LiteLLM supply chain compromise (March 2026): PyPI versions 1.82.7 and 1.82.8 were backdoored. Audit installed versions and upgrade to a clean release (LiteLLM Security Blog).
  • LiteLLM Docker tag deprecation: The main-stable tag is deprecated as of June 30, 2026 — migrate to :latest (LiteLLM Docs).

The concept of “agentic inequality” — disparities in power, opportunity, and outcomes arising from unequal access to AI agents across availability, quality, and quantity — gives this a name (arXiv: Agentic Inequality). The efficiency calculation that makes cost optimization look like a solution is the same calculation that renders invisible the costs it exports.

The Engineers Are Not Wrong

The steelman of aggressive cost optimization deserves to be stated fully, because it is genuinely persuasive. AI services that cannot control inference costs will price themselves out of the markets they claim to serve. A Batch API discount of fifty percent on inference makes AI products viable for organizations that could not otherwise afford them. Prompt caching reductions of up to ninety percent on cached input tokens mean that a small nonprofit running a legal aid chatbot pays a fraction of what it would without caching (Anthropic Docs). Open-source models like DeepSeek removed financial and technical barriers for underserved markets that proprietary API pricing had effectively excluded.

In this framing, LLM Observability tools that enable teams to understand exactly where tokens are spent — and eliminate the wasteful ones — are not instruments of inequality. They are what makes the difference between an AI service that reaches millions and one that serves thousands. The routing system that sends the majority of requests to a cheaper model and a smaller fraction to a premium one makes that premium fraction possible precisely because it conserved on the majority. Restricting cost optimization means restricting access. Who would that serve?

The Precise Point Where This Argument Fails

The defense rests on an implicit assumption: that quality preservation across the routing system as a whole is equivalent to quality preservation per user. It is not.

When a routing framework preserves most of premium quality at the system level, that measure aggregates across all requests and all users. It tells us nothing about whether the requests in the degraded fraction are randomly distributed. They almost certainly are not — and the research on routing fairness is explicit about the gap. Inference-time fairness routing for text Mixture-of-Experts models “remains uncharted,” with fairness-constrained training approaches deemed computationally prohibitive at scale (MoE Fairness paper, arXiv 2603.27141). The routing decisions that determine which requests reach stronger models are not designed with demographic equity as a constraint. They are designed with cost as the constraint.

This is not a technical oversight. It is a design choice — one made without the participation of the users who will bear its consequences. The EU AI Act’s full applicability in August 2026 does include non-discrimination requirements for high-risk AI deployments, though whether those requirements extend to general-purpose model tiering and routing systems remains genuinely contested (arXiv: EU AI Act fairness). The regulatory direction is clear. The current enforcement reach is bounded. And in the gap between direction and reach, the routing frameworks continue to operate without fairness constraints.

The Policy Nobody Voted For

Thesis: Aggressive LLM cost optimization is not neutral engineering but a series of policy choices about which users deserve quality — choices made without democratic mandate, surfaced through architecture rather than law, and distributed in ways that systematically disadvantage those with the least capacity to compensate.

The pricing spread between economy and premium model tiers — a 10x gap between the cheapest and most capable models in the current Claude family (Anthropic Docs) — is not a pricing decision in isolation. It is a decision about who can access the full capability of a technology that is increasingly consequential for economic and civic participation. When routing systems allocate the cheapest compute to users who cannot purchase their way into better treatment, they do not merely reflect existing inequality. They reproduce it, at inference speeds, across millions of interactions, without ever naming what they are doing.

The question is not whether cost optimization should exist. Of course the economics of AI infrastructure must be managed — the alternative is services that serve nobody. The question is whether cost optimization should be designed without equity constraints, and who decided, without any public process, that it could.

Where I Could Be Wrong

This thesis has vulnerabilities worth naming honestly. If routing systems were designed with fairness constraints that demonstrably counteracted socioeconomic bias — routing toward premium models precisely when signals indicate user disadvantage rather than advantage — then the mechanism I am describing would be interrupted. The research does not yet show this happening at scale, but the architecture would permit it, and some recent routing approaches have begun treating allocation as a multi-objective problem.

If the quality gap between frontier and economy model tiers continued to close at its current pace, the stakes of routing assignment would diminish. A world in which today’s economy-tier models approach the performance of current mid-tier models would make the tiering argument less urgent — though not moot, because the gap at the frontier would simply move upward, and the differential would persist in a different form.

And if the populations most affected by differential routing were found to benefit disproportionately from broader AI access — if Level 1 access gains reliably outweighed Level 2 and 3 quality losses — the balance of the argument would shift. I do not believe the current evidence supports that conclusion, but I hold the possibility seriously, because an argument for equity that ignores the genuine benefits of access would be its own form of bad faith.

The Question That Remains

The efficiency gains of LLM cost optimization are real, and they reach people who would otherwise have no access to capable AI at all. The inequalities those same optimizations produce are also real, and they land hardest on the people already carrying the most weight. The question the industry has not asked — not seriously, not in public — is whether infrastructure this consequential should be governed by cost alone, or whether equity deserves a place in the routing objective.

That is not a technical question. It is a political one. And the longer we pretend the engineering is neutral, the longer we defer asking who is authorized to make that choice — and who answers for the consequences when they get it wrong.

Ethically, Alan.

AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors