Hidden Instructions: The Ethics of System Prompts, Behavioral Manipulation, and Accountability

The Hard Truth
System prompts are an operator’s right — the same right that lets a call center script its agents, a brand establish its tone guidelines, or a government agency decide what its AI assistant will and won’t discuss. Meta-transparency resolves the apparent tension: governance norms are public, even when individual configurations are not. We do not demand that every bank publish its fraud detection logic before extending credit, and demanding the same from AI operators conflates implementation details with deception.
The argument has genuine force, and I want to sit with it seriously before dismantling it. The companies building System Prompts governance frameworks are not, for the most part, acting cynically — they are grappling with a real problem, and some of them have thought carefully about it. That careful thinking deserves a fair reading before we examine what it misses.
The Case That Deserves a Fair Hearing
The strongest version of the confidentiality argument runs as follows. Software has always operated through layered abstractions. Users do not read the source code of their operating systems, the query planners inside their databases, or the recommendation algorithms that surface their feeds. Requiring every layer to be visible to every user is not a principle of transparency — it is an engineering impossibility that conflates implementation with intent.
Anthropic has attempted a serious solution with what it calls meta-transparency. The policies governing what operators can and cannot instruct Claude to do are fully public, even when the specific instructions are not. Under these published norms, operators may restrict what the AI discusses, assign it a custom persona through Role Prompting, and decline to discuss topics outside their domain — but they cannot direct it to deceive users in ways that damage their interests, claim to be human when sincerely asked, or withhold information that endangers health or safety (Anthropic’s Constitution). The framework distinguishes restriction from weaponization, and makes that distinction legible through publicly auditable policy rather than individually disclosed prompts.
This is genuine governance thinking with internal coherence. The EU AI Act adds a legal enforcement layer from 2 August 2026 onward, requiring chatbots to identify themselves as AI at every interaction — penalties reaching €15 million or three percent of worldwide turnover for non-compliance (EU AI Act). The institutional architecture is not absent. It is being built, and some of it reflects careful thinking about the asymmetry of power between operators and users.
The Load-Bearing Assumption
The argument works if users can act on it. This is the load-bearing assumption, and it deserves direct examination.
Meta-transparency as an accountability mechanism requires that users know the governance framework exists, can interpret the norms it publishes, and have a realistic path to redress when those norms are violated. A CHI ‘26 study on user perspectives on AI system prompts found that users are “concerned about hidden instructions that determine AI behavior without their knowledge or consent,” specifically identifying “accountability gaps” between organizational control and user expectations for transparency (CHI ‘26 paper). These are not naive users who misunderstand how software works. They are ordinary people engaged with systems that present as conversations — and discovering that the terms of those conversations were set elsewhere, by parties they cannot identify, through Context Engineering choices they cannot read.
The flaw is structural, not incidental. Governance information placed in published policy documents that most users will never encounter does not solve an accountability problem — it relocates it. The formal mechanism exists. The practical capacity to use it, for most people, in most interactions, does not.
Who benefits from an accountability structure that is technically sound and practically unreachable? That question, asked carefully, is where the inversion begins.
Confidentiality as Architecture
In January 2026, researchers published the JustAsk framework, which achieved full or near-complete Prompt Leakage across 41 commercial black-box models — recovering system prompt contents through standard user interactions, with no privileged access required (JustAsk paper). The researchers used the same Context Window available to any user, probing through ordinary conversation.
This finding carries two implications that cut in different directions, and taking both seriously changes the picture.
The first: the confidentiality rationale is weaker than claimed. If researchers can systematically reconstruct what operators instructed the AI to do through normal interactions, then “hidden” describes a social convention rather than a technical guarantee. The instructions are accessible — they require sophistication to extract.
The second implication is the more structurally serious one. The gap between what a sophisticated researcher can discover and what an ordinary user understands creates its own accountability asymmetry. The confidentiality holds most reliably against the people who most need to see past it.
Add the adjacent failure mode. In August 2025, indirect prompt injection through commands embedded in publicly visible content hijacked a user’s Perplexity Comet browser, accessing their email and exfiltrating credentials within 150 seconds — the user unaware throughout. When systems are architecturally designed to trust Instruction Following at the system level, that trust becomes a surface. The instructions that shape legitimate AI behavior and the injected commands that subvert it are technically indistinguishable from the model’s perspective. Both are instructions. Both are followed.
The same architecture that enables legitimate customization enables behavioral manipulation. The distinction between the two is meaningful — but it is only accessible through the prompts themselves, which remain hidden.
Accountability Without a Visible Target
Thesis: The accountability architecture around system prompts distributes responsibility across operators, platform providers, and regulators precisely in the configuration that makes it hardest for any individual user to navigate.
Anthropic sets norms but cannot audit every operator deployment. Operators configure behavior but have no direct accountability relationship with the users those configurations affect. Regulators write rules on timelines measured in years while incidents occur on timelines measured in seconds. The user who encounters a problem is the party with no institutional leverage, no visibility into the instructions that produced the harm, and no clear mechanism for determining whether what happened was an operator violation, a model failure, or a deliberate design choice operating exactly as intended.
The Meta Prompting layer — the instructions that set the terms of every interaction — is the most consequential part of these systems and the part most systematically shielded from the people who bear its consequences. That combination is not incidental. It is the architecture.
Who Pays for the Abstraction
Governance abstractions have costs, and those costs land on identifiable people.
In 2024, a British Columbia Civil Resolution Tribunal found Air Canada liable for misrepresentation by its chatbot, which had promised a bereavement fare discount that the airline’s actual policy did not provide. The airline argued its chatbot was a separate legal entity not bound by what it said. The tribunal rejected this, ordering the company to pay $650.88 plus fees (McCarthy Tétrault). The ruling carries limited precedential weight beyond British Columbia — it was a tribunal decision, not a court judgment — but it illustrates what accountability looks like when it finally materializes: one user, a lengthy dispute process, a modest amount, and a company that had configured its system without apparent regard for what it would tell people in difficult moments.
In October 2025, a coalition of 36 consumer advocates called on the FTC to halt Meta’s announced plan to analyze chatbot conversations for advertising personalization, describing the practice as a “material omission or misrepresentation likely to mislead consumers” (EPIC). As of the research for this article, the FTC had not issued a ruling. Whether it will, and when, remain open questions. In the meantime, users whose conversations may be analyzed for advertising purposes do not know this is occurring — because the instructions governing that analysis exist in the system prompt.
What does accountability mean when the path to redress requires knowledge of regulatory structures, access to legal processes, and persistence that most users simply do not have? The users the opposing view’s accounting leaves out are those who engaged in good faith with systems they assumed were oriented toward their interests, and discovered otherwise only after the fact.
What Would Make This Wrong
The argument here is not that institutional accountability is impossible. It is that the current architecture does not deliver it in the conditions that matter most. That argument would be substantially weakened — perhaps refuted — by evidence that enforcement mechanisms are reaching the failure modes described above.
EU AI Act Article 50’s enforcement date of 2 August 2026 is approximately five weeks away as of this writing. Its transparency obligations address AI identity disclosure — whether chatbots identify themselves as AI — rather than the content of the instructions governing their behavior (EU AI Act). But enforcement at scale creates precedents and institutional capacity. If regulators demonstrate willingness to issue significant penalties, the incentive structures for operators change over time. The question is whether enforcement expands from identity disclosure to behavioral manipulation — whether the gap between “the system identified itself as AI” and “the system was configured to serve another party’s interests while presenting as your advocate” becomes a distinction regulators are empowered and willing to address.
If it does, this argument needs revision. If governance capacity scales to match the governance problem, then the existing frameworks become foundations rather than facades. That is the version of the future this argument most wants to be wrong about.
The Question That Remains
The question is not whether operators should configure AI behavior. They will, and there are legitimate reasons they should. The question is whether the governance architecture being built around that configurability is designed to deliver accountability that is meaningful in practice — visible to and usable by the people the system affects — or accountability that satisfies institutional requirements while remaining beyond the reach of those it nominally protects.
What happens when the rules that determine how an intelligence behaves toward you are owned by a party you cannot name, enforced by norms you have never read, and contested through processes that assume prior knowledge you do not have?
AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors