Liability Without Transparency: Ethical Risks of Domain-Specific Prompting in Regulated Industries

The Hard Truth
More than 700 court cases now involve AI-generated content — hallucinated citations, fabricated precedents, unsupported recommendations (LexisNexis/Bloomberg Law, via SQ Magazine). The number circulates as both an indictment of AI and evidence that our institutions are adapting. Neither interpretation asks the question that matters most: what were those systems told to do, and does anyone know?
The courtroom is one of the few places where accountability eventually surfaces. When a lawyer submits a brief with fabricated citations, a judge notices. What these moments of reckoning share is that they arrive downstream — errors have already been made, harm is already in motion, and only then does the system ask what went wrong. What it almost never asks is what went in.
Seven Hundred Cases and What They Appear to Settle
Domain-Specific Prompting is what separates a general-purpose AI assistant from the system a law firm uses to analyze contract clauses, a hospital integrates into triage workflows, or a financial adviser uses to generate compliance summaries. This specialization is achieved through Prompt Engineering — the deliberate construction of instructions, constraints, and context that shape how a model reasons within a particular domain. The outputs are domain-specific, consequential, and increasingly professional in register. The inputs — the decisions about what the system should know, how it should prioritize, what kinds of uncertainty it should flag — are rarely visible to the people whose professional futures those outputs affect.
The 700+ court cases involving AI-generated hallucinations and fabricated citations (LexisNexis/Bloomberg Law, via SQ Magazine) appear to settle a proposition: AI in high-stakes professional settings is risky. But the question they appear to settle is not the interesting one. The more interesting question is why these cases are arriving in court at all, and what that pattern reveals about a governance layer nobody examined before the harm occurred.
Two Readings of the Same Count
There is an optimistic reading of 700 cases. Courts are catching errors. Judges are sanctioning attorneys who submit AI-generated fiction as legal precedent. The adversarial structure of the legal system — built over centuries to surface contested claims — is adapting to a new category of unreliable assertion. From this perspective, the number is evidence that correction remains possible; that the institutions designed to allocate responsibility for harm are still functioning.
The critical reading is darker. Seven hundred is not the total. It is the visible fraction: cases where someone noticed, where the error was significant enough to be challenged, where the affected party had both the resources and the standing to object. Secondary research summaries suggest that citation fabrication rates on specific legal queries may reach substantial levels — though these figures derive from indirect sources and should be treated as directional, not definitive. The gap between errors made and errors caught is wide, and it grows in proportion to adoption. Every unchallenged AI output that shaped a legal argument, a clinical recommendation, or a credit assessment that was not contested — because the affected party accepted it, did not know AI was involved, or could not afford to object — exists outside this count.
Both readings agree that harm is occurring. They disagree about scale. But they share something more important than their divergence.
The Assumption Neither Reading Questions
The assumption is that when harm is traced to an AI system, the instruction layer is available for inspection — that there exists, somewhere, a legible record of what the model was told. What System Prompts governed its behavior. What constraints shaped its outputs. What domain-specific rules determined how it framed risk, and how confidently.
In most current implementations, that assumption is false.
Knowledge Injection — the practice of encoding domain-specific expertise, regulatory frameworks, and institutional norms directly into a model’s operating context — is typically proprietary. A legal AI tool’s instructions about how to handle jurisdiction-specific precedent, how to qualify ambiguity, and how to weight competing interpretations live in vendor-controlled configurations that neither the supervising attorney nor the affected client can access. When a clinical decision-support system produces a treatment recommendation, the physician sees the recommendation. The Context Window that produced it — the hierarchy of instructions that weighted some clinical evidence over others, the Role Prompting that assigned the model a clinical persona with specific epistemic commitments — none of this is visible at the point where the professional decision is made.
The professional is being held accountable for the consequences of a system she cannot read.
When the Audit Trail Leads Nowhere
The American Bar Association’s Formal Opinion 512, released July 29, 2024 — the first ethics guidance the ABA has issued on generative AI for lawyers — makes the duty of competence explicit: attorneys must understand AI’s risks and benefits before using these tools in practice (ABA). It addresses confidentiality, candor toward tribunals, and supervisory responsibility. What it cannot address, structurally, is transparency into prompt configurations the attorney did not author and cannot inspect. By early 2026, more than 35 state bar associations had issued supplemental guidance. All of them share the same structural gap: they describe professional obligations without the mechanisms to fulfill them when the underlying instruction layer is invisible.
In medicine, the FDA’s revised Clinical Decision Support Software guidance, issued January 6, 2026, loosened oversight of AI tools that fall outside strict device classification and explicitly placed liability on the clinician when AI-generated recommendations are involved (Arnold & Porter). The Instruction Following behavior that makes the model confident in its suggestions, the Context Engineering choices that determined which clinical guidelines were weighted most heavily — these are design decisions with clinical consequences for which the physician is now responsible, without any mechanism to inspect them.
Finance presents a structurally similar arrangement. The 2026 FINRA Annual Regulatory Oversight Report confirmed that existing supervision and recordkeeping rules apply to all AI-generated content (FINRA). Yet a substantial proportion of investment adviser firms using AI had no formal testing or validation processes for their AI outputs at the time. The Multi-Turn Prompt Design underlying a financial AI system — the accumulating sequence of instructions that builds context across a client interaction — is as much a compliance artifact as the trade it influences. It is rarely treated as one. The firm bears the regulatory exposure while the vendor controls the specification.
The EU AI Act classifies AI assisting judicial authorities, AI embedded in medical devices, and AI used for creditworthiness assessment as high-risk under Annex III, subject to full compliance obligations (EU AI Act Annex III). A proposed “Digital AI Omnibus” amendment may defer some of those deadlines beyond August 2026, though the final status remains pending (DLA Piper). But even when the Act’s requirements fully apply, what constitutes adequate documentation for a high-risk system’s instruction layer has not been established. The classification is clear; the accountability mechanism is not.
The Governance Layer Nobody Is Governing
Thesis: The liability question in regulated industries is not whether AI-generated errors cause harm — they demonstrably do — but whether we will allow the instruction layer that shapes those errors to remain permanently invisible, allocating accountability to practitioners who cannot see what they are accountable for.
When prompt engineering encodes institutional judgment — how a clinical AI frames diagnostic uncertainty, how a legal AI qualifies contractual risk, how a financial AI assesses the plausibility of a borrower’s claims — that judgment becomes a de facto policy. It acquires the force of a rule without any of governance’s obligations: no public debate, no audit trail, no principled mechanism to challenge it.
Multimodal Prompting extends this concern into new territory: in 2026, the attack surface for domain-specific AI systems includes not just text instructions but image-embedded and document-embedded injection vectors, widening the gap between what a system is actually doing and what its operators believe it is doing (OWASP Gen AI). The sophistication of context-engineering practice has outpaced the governance of context-engineering decisions by a distance that is not narrowing.
What No Court Has Yet Counted
The 700 cases count the errors that surfaced. They do not count the decisions that were subtly distorted rather than obviously wrong — the risk assessment that was systematically conservative for one demographic because of how domain knowledge was encoded; the legal opinion that accurately cited existing law but framed the client’s options through a prior probability distribution a vendor decided to emphasize; the clinical recommendation that was technically defensible but shaped by weights the prescribing physician could not see and did not know to question.
These are not errors in the legal sense. They are something closer to invisible policy — consequential choices encoded into systems that appear to be tools but function as governance, with none of governance’s obligations to be visible, challengeable, or reversible. The context-engineering literature increasingly treats prompt design as an engineering discipline with its own quality standards. It has not yet produced a tradition of treating prompt design as a governance artifact with corresponding accountability obligations.
What Would Change My View
This argument depends on a claim about invisibility, and that claim is most vulnerable at its institutional edge. If domain-specific prompt configurations were routinely documented, version-controlled, and made available to supervising practitioners under appropriate confidentiality agreements, the governance gap I describe would be meaningfully reduced. There is emerging work on prompt auditability — configurations that surface their own assumptions to users rather than concealing them. If that approach reached regulated industries at scale, the liability-without-transparency problem would not disappear, but it would become tractable.
I would also revisit this position if the EU AI Act’s Annex III requirements were interpreted by regulators to include prompt documentation as mandatory technical documentation for high-risk AI. Regulations, when enforced with sufficient specificity, can create accountability obligations that professional ethics alone have not managed to produce.
The Question That Remains
The 700 cases represent accountability arriving after the fact. The underlying structure — where the instruction layer shaping high-stakes decisions is proprietary, undocumented, and outside the professional visibility of those now held responsible for its outputs — has not changed. It is accumulating.
What changes first: the volume of harm that makes invisibility legally untenable, or the governance framework that treats prompt documentation as a professional obligation before the next wave of cases arrives?
AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors