Intelligence (AI) The Complete Chain of AI Evidence and Accountability Claim → source → provenance chain → version/model → process → human review → decision → impact → audit trail → correction/appeal → responsibility Fluent answer ≠ evidence; automation ≠ removal of accountability; “the model said so” ≠ adequate reason; decision without an audit trail = difficult to re-examine; acknowledgement of error + correction + appeal = accountable system Problem In the age of artificial intelligence, the crisis of knowledge is not merely a crisis of wrong answers; it is changing the location of evidence, the limits of authorship, the source of decisions and the structure of responsibility. Earlier, a reader could usually identify an author, institution, document or direct source. In an AI-generated answer, the sentences may be smooth, concise and confident, yet it is not directly visible which source supports each claim, what training material influenced it, what retrieval or tool use occurred, or what part is inference. This widens the distance between “receiving an answer” and “acquiring knowledge.” If users treat AI output as evidence by default, linguistic skill can take the place of truth. Fluency, elegant structure, technical vocabulary and citation-like form become psychological signals of credibility; epistemic reliability, however, requires sources, method, consistency, reproducibility, counter-evidence and a history of revision. AI decision systems, moreover, do not merely generate knowledge; they can affect human opportunities in recruitment, credit, insurance, education, medicine, security, content moderation, evaluation, fraud detection, resource allocation and administration. In these fields an error does not mean merely a “wrong answer”; it may affect employment, reputation, money, services, liberty or rights. Here a familiar accountability problem arises: a developer may say that the user misused the system; the deployer may say the vendor built the model; a manager may say the algorithm only made a recommendation; an operator may say the system auto- approved the case; and the institution may say that the procedure was standardised. If everyone participates in the causal chain but no one is finally answerable, a responsibility gap appears. The complexity of an AI system does not reduce responsibility; it increases the need for documentation. An opaque model may not always yield a simple account of its exact internal causal path, but an institution can still state the model and version, the inputs, policy, thresholds, whether human override existed, how validation was performed, known limitations, who reviewed the decision, and where an appeal can be made. Evidence is also time-sensitive. A model update, change in the retrieval corpus, system prompt, safety policy, tool access, temperature, memory or user context may change the answer. Thus “AI said this” without a date, model/version, context and source snapshot is an incomplete citation. A different answer to the same question later raises a problem of reproducibility. AI-generated citations pose a particular danger. A real author or book may be paired with a wrong title, a correct title with a wrong page, or an entirely fabricated reference may appear. Citation-like form is therefore not itself a citation. Verification in the original publication, bibliographic record, DOI/ISBN record or a trustworthy catalogue is necessary. Accountability is not merely the search for someone to blame after an error. Responsible design begins earlier, with risk anticipation, data governance, testing, documentation, access control, human review, monitoring, incident reporting, correction and redress. Punishment comes only at the last stage; prevention, transparency and institutional learning are also parts of accountability. There is an opposite danger as well: dismissing every use of AI as unreliable. Machine-assisted search, translation, summarisation, coding, pattern detection and decision support can be useful in many settings. The proper response is neither blind trust nor blanket rejection, but task-specific standards of evidence and impact-specific oversight. The purpose of this chapter is not to produce a generic list of AI ethics principles, but to clarify the philosophical relation between evidence and accountability. A claim without evidence provides a weak basis for decision; a decision that is not traceable is difficult to review; review without appeal can be unjust; and an institution that does not correct errors cannot develop stable knowledge. Core Proposition The central proposition of this chapter is: in the AI age, understand evidence not merely as the “content of an answer” but as a chain of provenance; and construct accountability not as “blame attached to one individual” but as role-specific duties together with final answerability. Verifiable sources, versioned processes, human authorisation, audit trails, correction and appeal are elements of one accountable knowledge system. First principle—AI output may be prima facie information, but it is not automatically evidence. Verification may be light for ordinary brainstorming; in high-stakes medicine, law, finance, historical quotation, academic references or public policy, independent checking of the original source should be treated as mandatory. Second principle—the burden of evidence varies with the claim. “This book is well known” and “this sentence occurs on page X” do not require the same verification. Numbers, dates, quotations, legal provisions, drug dosages, scientific results, attribution and archival claims demand high source-specificity. Third principle—provenance means the chain of origin. Which model/version, prompt/context, retrieved documents, tools, human edits and time produced the answer? Such metadata makes a claim more reproducible. Provenance is not a complete guarantee of truth, but it is the basis of audit. Fourth principle—explanation and evidence are distinct. A system may generate a plausible reason after a decision even when that reason does not reflect the actual causal process. A natural-language answer to “why?” should therefore be checked against decision logs, feature records, rules, thresholds, sources or causal tests. A post-hoc story is not evidence. Fifth principle—“human in the loop” is not merely a label. If a reviewer mechanically approves thousands of recommendations, human presence does not amount to genuine oversight. The reviewer needs time, authority, competence, relevant information, power to override and a documented reason. Sixth principle—automation bias is an enemy of accountability. If a person treats machine recommendations as objective and relinquishes responsibility for judgement, errors may be amplified. The opposite, algorithm aversion, is also unreasonable: one observed error does not justify discarding a useful system. Policy should be based on comparative human-machine performance and context. Seventh principle—responsibility may be divided, but it must not be dissolved. Data providers, model developers, integrators, deployers, domain professionals, managers and institutions may have different duties; do not make everyone responsible in a way that leaves no one responsible. Duties at each stage should be explicit, while final accountability to the affected person should rest with a clearly identified institution. Eighth principle—appeal is also an epistemic mechanism. When an affected person challenges a decision, they are not only seeking justice; they may supply new evidence of system error, loss of context, identity mismatch, outdated data or proxy failure. An appeal system is a source of institutional learning. गजेन्द्र ठाकु रक समानान्तर दर्शन — खण्ड २ Ninth principle—a history of correction is part of knowledge. Silently replacing an erroneous output is insufficient once earlier decisions have already had effects. Significant corrections should record what changed, why it changed, which earlier cases may be affected, and whether re-examination is required. Tenth principle—in high-impact settings, “cannot determine” is a valid output. When uncertainty is high, data are missing, evidence conflicts, or a case is out of distribution, forcing a binary answer is dangerous. Abstention, escalation and human review are signs of reliable design, not weakness. Eleventh principle—model confidence and truth are not identical. A probability-like score may be poorly calibrated, and the tone of certainty in natural language is not itself technical confidence. User-facing language about certainty should correspond to the strength of the evidence. Twelfth principle—the final question of accountability is always normative: “Who had the authority to make this decision, who had a duty of care, who had the capacity to correct the consequences, and before whom can the affected person demand an answer?” Technical architecture informs these questions; it does not replace them. Principal Arguments Argument 1—source-free fluency creates an epistemic shortcut. A language model may be highly skilled at producing coherent sentences, but coherence is not a sufficient condition for truth-tracking. If users take prose quality as a proxy for truth, hallucination becomes especially persuasive. High-value claims should therefore be linked to sources. Argument 2—retrieval does not automatically improve evidential status. A retrieved document may be wrong, outdated, biased, incomplete or mismatched to the query. In RAG-like systems, having a source is useful, but source quality, passage relevance, fidelity of quotation and accuracy of synthesis must still be checked separately. Argument 3—a correct source can still support an incorrect inference. Even with an authentic document, a model may omit context, treat correlation as causation, ignore exceptions, or flatten disagreement among sources. Citation does not replace an audit of reasoning. Argument 4—opacity of training data makes attribution difficult. Even when generated text is not an exact quotation, its style, facts and ideas may statistically mix many sources. Ethical analysis of authorship therefore needs to distinguish memorisation, transformation, retrieval, public-domain status, licence and attribution policy. Argument 5—a decision score is not a moral decision. If a risk score is 0.72 and the denial threshold is 0.70, who chose that threshold? How were the costs of false positives and false negatives weighed? What is the effect on protected groups? The score is a technical output; the threshold is a normative policy. Argument 6—proxy selection carries power. Abstract qualities such as merit, risk, fraud, quality, engagement and toxicity cannot usually be measured directly; features and proxies are chosen. The definition of a proxy embeds social values. The classificatory discipline of Chapter 86 here becomes a question of accountability. Argument 7—the distribution of error may matter more than average accuracy. An overall accuracy of 95 per cent can conceal a rate of only 65 per cent for a small group. Aggregate metrics are incomplete without disaggregated testing, subgroup uncertainty, sample size and context-specific harm. Argument 8—non-determinism increases the need for audit. The same prompt may produce different outputs, and system updates may change earlier behaviour. Incident investigation should therefore capture the prompt, model/version, settings, retrieved sources, tool results, timestamp and final human action. Argument 9—accountability is hollow without genuine freedom to override. If employees are punished for changing an algorithmic recommendation, or the interface strongly pushes a default choice, formal human control is not real. Governance must examine interfaces and incentives as well as formal rules. Argument 10—explainability is audience-specific. Feature attribution may help an engineer; a doctor needs clinical rationale; an applicant needs an understandable reason for an adverse decision; a regulator needs a process record. One technical dashboard cannot satisfy every form of accountability. Argument 11—privacy and transparency require balance. Publishing all training data or user logs may violate privacy, trade secrets or security. Accountable transparency may instead combine role-based access, independent audit, aggregated reporting, confidential regulatory review and explanation to affected persons. Argument 12—security may justify limited disclosure, but it is not a universal excuse. Exact exploit details may properly remain restricted, yet system category, risks, incident counts, mitigation status, appeal processes and oversight structures can still be disclosed. Argument 13—liability depends not only on causation but also on duty. Some technical contributors may be causally remote, yet a deployer that knowingly places a system in a high-stakes context has a stronger validation duty. In general, responsibility tends to increase with proximity to harm and degree of control. Argument 14—blaming an individual can conceal institutional design. An operator may click wrongly because the interface is confusing, workloads are impossible, alerts are excessive, training is inadequate and override policy is unclear. “Human error” is not a final explanation; the sociotechnical system must be audited. Argument 15—documentation is not a static PDF but a lifecycle record. Model cards, data documentation, evaluation reports, change logs, incident logs, access policies, monitoring results, user complaints and remediation should remain connected over time. Argument 16—accountability is incomplete without correction. “We regret the error” is insufficient if a wrong decision remains in the user’s record. Correction must reach downstream databases, notifications, affected decisions and financial or administrative consequences. Argument 17—the responsibility gap is not solved merely by declaring the machine a moral agent. In most present institutional settings, legal and organisational authority remains with human institutions. However autonomous a machine output may appear, the decision to deploy it is part of human governance. Argument 18—sometimes the system’s best action is escalation. In cases of high ambiguity, conflicting evidence, vulnerable persons, rare cases, novel legal issues or safety-critical contexts, “manual review required” may be a sign of reliable design. Maximising automation coverage is not always a proper goal. Pūrvapakṣa First Pūrvapakṣa: “AI is more accurate than humans on average, so humans should be removed everywhere.” Average performance does not show context, subgroup performance, rare cases or the cost of accountability failure. A more accurate system may be useful, but transfer of authority requires examination of failure modes, appeal, domain shift and severity of harm. Second Pūrvapakṣa: “AI is only a tool; responsibility belongs entirely to the user.” The design of a tool, its defaults, warnings, capability claims, deployment constraints and foreseeable misuse can create duties for developers and deployers. Analogies with simple tools such as hammers do not always apply to adaptive opaque systems. Third Pūrvapakṣa: “The model made the decision autonomously, so no human is responsible.” Humans and institutions determine the environment, objectives, data access, thresholds, deployment and authority within which autonomous action occurs. Operational autonomy is a description, not an argument for an accountability vacuum. गजेन्द्र ठाकु र Fourth Pūrvapakṣa: “A source is provided, so the answer is proven.” The existence of a source is only the first stage. One must still examine the relevant passage, exact quotation, authority of the source, date, context, conflicting sources and validity of inference. Fifth Pūrvapakṣa: “Explainable AI solves every problem.” Fidelity of explanation, comprehensibility, audience, manipulability and causal relevance remain open questions. A beautiful explanation cannot legitimise an unjust policy. Sixth Pūrvapakṣa: “The system is proprietary, so transparency is impossible.” Trade secrets may limit full public disclosure, but confidential independent audit, regulator access, performance reporting, reasons for adverse action and incident disclosure remain possible. Seventh Pūrvapakṣa: “There is a human reviewer, so the system is accountable.” If the reviewer merely rubber-stamps decisions, lacks time, lacks override authority or lacks training to understand the model’s logic, formal review is only a symbol of accountability. Eighth Pūrvapakṣa: “Remove protected attributes and bias disappears.” Proxy features can reconstruct signals of protected attributes, while definitions of fairness—outcomes, error rates, opportunity, calibration and others—can conflict. Attribute blindness does not automatically produce justice. Ninth Pūrvapakṣa: “A very large model is inherently robust.” Scale may increase capability, but hallucination, rare failures, prompt sensitivity, emergent misuse and domain-specific errors do not automatically disappear. Testing should be based on task and harm, not model size. Tenth Pūrvapakṣa: “AI decisions are objective, whereas humans are subjective.” Training labels, loss functions, feature selection, thresholds, data collection and policy objectives all embody human choices. Machine consistency may be one element of objectivity, but it is not value-neutrality. Eleventh Pūrvapakṣa: “Appeal reduces efficiency and destroys the benefit of automation.” Appeal is a governance cost, but high- stakes errors may cost more. Targeted escalation, sampling, risk tiers and efficient review design can create a balance. Twelfth Pūrvapakṣa: “Errors are inevitable, so accountability is unfair.” Accountability need not demand zero error; it can be grounded in reasonable care, management of known risks, truthful capability claims, monitoring, prompt correction and redress. Uttarapakṣa First rule of the Uttarapakṣa—classify the claim. General explanation, exact fact, number, quotation, diagnosis, legal interpretation, prediction, recommendation, ranking and eligibility decision require different evidential thresholds. The system interface should make these differences visible to the user. Second rule—create a source ladder. Primary documents and original data have the highest priority; reputable secondary analysis comes next; tertiary summaries are auxiliary; uncited AI prose should be treated only as navigation or brainstorming. The more important the claim, the higher the required source should be on this ladder. Third rule—verify citations both automatically and by human review. Machines can check whether a URL or DOI exists, whether a quoted string matches, publication metadata and freshness of date; humans or specialists must still examine semantic relevance, context and contested interpretation. Fourth rule—maintain a minimum provenance log: timestamp, model/version, system configuration, user prompt or case identifier, retrieved source IDs, tool actions, decision rule, human reviewer and final action. Privacy-sensitive fields should remain access-controlled. Fifth rule—for high-stakes decisions, record a structured “reason for action”: which evidence was considered, which evidence was excluded, the key rule or threshold, uncertainty, human reviewer and available appeal. “The system determined this” is not a reason. Sixth rule—record overrides in both directions. If the reviewer changes the algorithm’s recommendation, record why; if the recommendation is followed, especially in high-risk cases, record an affirmative review rather than blind acceptance. These records allow later learning. Seventh rule—create incident-severity tiers. A harmless formatting error and denial of a medical service should not trigger the same process. Escalation timelines should reflect severity, reach, reversibility, affected population and probability of recurrence. Eighth rule—maintain rollback readiness. If a model update produces harmful results, there should be a version registry, change log and capacity to return to an earlier validated version. “Continuous improvement” without rollback can become an uncontrolled experiment. Ninth rule—define the scope of independent audit clearly. Audit should examine data sampling, performance, subgroup errors, security, privacy, documentation, human workflow, complaint records and governance authority—not only model benchmarks. Tenth rule—ensure intelligibility for the affected person. An explanation may be technically correct but useless for redress if the person cannot understand it. A simple explanation should be available to the individual, with a more detailed technical record available at another level. Eleventh rule—make uncertainty disclosure actionable. Do not merely write “low confidence”; define what the system does when confidence is low—abstain, request more evidence, consult a second model, seek expert review or place a temporary hold. Twelfth rule—monitor outcomes, not merely system operation. After deployment, measure not only uptime and latency but errors, complaints, overrides, subgroup performance, drift, appeal reversals, harm incidents and correction latency. Thirteenth rule—public claims must correspond to audited evidence. Marketing phrases such as “human-level,” “bias-free,” “fully explainable” or “safe” can mislead without defined benchmarks and material qualifications. Accountability begins with truthful communication of capability. Fourteenth rule—identify a final accountable office. Do not force users to wander among vendor, deployer and operator. One institution or office should receive complaints, coordinate investigation, deliver correction and then allocate internal liability. Indian Dialogue Indian epistemology is not a software manual for the AI age, but it offers deep resources for disciplining knowledge claims. Different traditions recognise perception, inference, testimony, comparison, postulation and non-apprehension in different ways; the central lesson is that “something was said” and “knowledge was established” are not the same thing. The Nyāya tradition discusses pramāṇa as producing veridical cognition and analyses defects and pseudo-evidence to understand epistemic failure. If AI output functions in a testimony-like way, modern analogues of competence, reliability, intention, defect and corroboration arise. It is not automatically justified to treat a machine as an āpta, an authoritative knower. The analogy with verbal testimony should be handled carefully. Classical theories of testimony involve complex questions about speaker, sentence and meaning, whereas language-model text may be a statistical synthesis of many sources. “AI said this” is therefore not itself testimonial authority; a source-grounded chain is needed. The Nyāya concern with vyāpti in inference is illuminating for algorithmic prediction. A model generalises from relations between features and outcomes in historical data, but hidden conditions, domain shift, confounding or institutional change may break the relation. The discipline of testing vyāpti has a useful parallel in modern validation. The fine-grained Navya-Nyāya analysis of delimitation, relation, qualifier and locus prompts useful questions in classification: “Of what exactly is ‘risk’ a property, at what time, in what context, and within what limit?” A context-free label is philosophically incomplete. गजेन्द्र ठाकु रक समानान्तर दर्शन — खण्ड २ The distinction between doubt and determination is especially important. When evidence has equal force, data are incomplete or sources conflict, doubt is a legitimate epistemic state. Removing the pressure on an AI system to produce a definite sentence for every question accords with a Nyāya-like methodological discipline. The Mīmāṃsā debates over intrinsic and extrinsic validity can be fruitfully compared with trust design. Should output be accepted prima facie unless a defeater appears, or should external verification always be required? The answer may be domain- dependent: prima facie trust in low-stakes routine tasks, external checking for high-stakes specific claims. Buddhist epistemological traditions attend to the momentariness of cognition, error, conceptual construction and the limits of evidence. This is a useful reminder against treating an AI-generated category as the intrinsic nature of an object; the label may be a conceptual and operational construction rather than an ultimate reality. A proper use of Jain anekāntavāda does not mean “all views are equal,” but recognition of limits of standpoint. One model may describe economic risk, another clinical risk, a third fraud risk; each describes from a different standpoint. Decision-making therefore requires explicit specification of viewpoint. It would be inappropriate to impose the Gītā or Dharmaśāstric ideas of duty directly upon modern corporate liability, yet the question of role-relative duty remains useful: the developer’s duty, the domain professional’s duty and the administrator’s duty differ. Clear roles reduce diffusion of responsibility. The institutionalised form of Pūrvapakṣa and Uttarapakṣa in Indian philosophical debate can inspire modern audit culture. Do not record only the system owner’s claim; formally record the strongest objection, the response, remaining doubt and the reason for the final decision. The broader lesson is that Indian traditions of evidence do not perform a theatre of certainty; they distinguish instruments of knowledge, defects, doubt, defeat and validity. AI governance requires comparable subtlety. Mithila’s Parallel Perspective Mithila’s Nyāya–Navya-Nyāya tradition offers a particularly useful methodological model for AI evidence and accountability: state the proposition precisely, specify the reason, test examples and pervasion, search for hidden conditions, hear the opposing case, and write down the limits of the conclusion. Against black-box authority, this reason-giving culture is important. The Gangeśa-style mode of questioning—what kind of cognition is this, what caused it, what defects are possible, what generates doubt—can become an audit checklist for AI evaluation. Instead of saying “the model is accurate,” ask: for which task, population, time period, comparator, error class and confidence interval? A limited proposition is more evidentially responsible. The concept of upādhi offers a deep parallel for algorithmic fairness. A relation “feature X → risk” may appear true in training data while a hidden condition Y is actually responsible; when policy changes, the relation may break. Searching for hidden conditions is not merely statistical refinement but a philosophical responsibility. The logic of absence disciplines negative evidence. “No complaint was received” does not mean “no harm occurred” when the complaint mechanism is inaccessible. Absence becomes meaningful evidence only when conditions for detection are adequate. Governance dashboards should adopt this logic. Anuvyavasāya, or second-order cognition, has a parallel in institutional self-audit: the system makes decisions, while a meta- layer examines decision quality, uncertainty, overrides, complaints and corrections. First-order output alone is insufficient; second-order monitoring is required. The Mithila tradition of śāstrārtha, in giving the opponent a formal place, provides a cultural model for accountability. Objections from independent reviewers, affected communities, domain experts and developers should all be recorded. An internal compliance checklist alone is insufficient. Parallel Philosophy does not use classical terms here as decorative labels. Its purpose is to enrich modern AI governance with the rigour of evidence, doubt, hidden condition, absence, causation and counter-position; and in return to test the limits of older concepts through new technological cases. From this perspective follows a maxim of accountability: “Where there is impact from a decision, there must be an account of its reasons; where there is an account of reasons, there must be room for opposition; where there is error, there must be correction; where there is an affected person, there must be a door to appeal.” Contemporary Applications In research writing, AI use should follow an explicit workflow. Brainstorming, language polishing, search assistance, citation discovery and final factual claims are different stages. Every bibliographic reference should be verified against the original catalogue or publication, quotations matched to page-level sources, and AI-generated summaries never treated as primary evidence. In journalism, breaking news makes a model’s outdated knowledge especially risky. Newsroom AI output should be verified by a source desk before publication; generated quotations should be prohibited; image and audio provenance should be checked; and correction logs should be public. In education, teachers can require a “source trail” alongside AI-generated answers. Students should locate sources for factual claims, identify AI errors and compare conflicting evidence. This can teach epistemic literacy more effectively than a simple ban on AI. In university assessment, disciplinary decisions should not rest on a detector alone. Detector error, writing history, oral clarification, drafts, citations and process evidence should be considered together, with an appeal path for accusations. In medicine, clinical decision-support recommendations should not replace a doctor’s judgement in high-stakes diagnosis or treatment. The version of the source guideline, patient-specific contraindications, uncertainty, reasons for override and follow- up outcomes should be recorded. In law, every AI-generated case citation should be checked in the original reporter or database. Fabricated precedent can corrupt judicial process. A practitioner’s professional duty cannot be transferred to a vendor. In recruitment, an institution deploying a résumé-scoring tool should ensure job-relevance validation, subgroup error testing, accommodation, human review and candidate appeal. “The vendor’s system is proprietary” is not a reason to block accountability to applicants. In credit and insurance, an adverse decision requires a meaningful reason. Feature attribution is insufficient if the person cannot understand what can be corrected. Mechanisms for updating erroneous data and obtaining re-evaluation should exist. In public-benefit administration, automation errors can have disproportionate effects on vulnerable citizens. Instead of default denial, systems should provide uncertainty escalation, offline appeal, explanation in local languages and access to human caseworkers. In policing and security analytics, a risk score must not automatically substitute for reasonable suspicion or guilt. False positives, demographic skew, data provenance and feedback loops require judicial and institutional oversight. In facial recognition, confidence scores are context-sensitive. Poor lighting, ageing, demographic imbalance, database quality and look-alike risk all matter. High-impact decisions require multiple items of evidence; a single model match should not be decisive. In content moderation, a label such as “toxicity” depends on context, dialect, satire, quotation and reclamation. The severity of automated action should be proportionate, with appeal, human contextual review and evaluation by relevant language communities. गजेन्द्र ठाकु र In social recommendation systems, an engagement objective can amplify misinformation and anger. Platform accountability therefore extends beyond individual posts to ranking objectives, exposure patterns, feedback loops and the effectiveness of mitigation. For generative image and video systems, provenance metadata, watermarking or signatures where appropriate, editing history and disclosure of sources can help control misinformation. Because metadata can be removed, provenance is not the sole test of truth; forensic and contextual verification are also needed. In customer-service AI, decisions such as refunds, cancellations and account closures should have clearly defined authority limits. If a bot makes a false promise, the institution cannot evade contractual or consumer responsibility by saying “the AI made a mistake.” A transcript audit trail is useful. For coding assistants, generated code in security-sensitive contexts requires review, testing, dependency audit and attention to licence and provenance. Code that compiles is not thereby secure, lawful or well designed. In scientific discovery, a model may generate hypotheses, but this does not remove the need for experiment, data provenance, statistical validation and replication. Credit for AI-assisted discovery should distinguish human and machine contributions. In archives and digital humanities, OCR or AI transcription should remain linked to the original image, uncertainty should be marked, and silent normalisation avoided. A “clean” historical text without a record of variants can distort the source. In language technology, AI output for under-resourced languages requires community review. Dominant-language patterns may distort local grammar and meaning. Provenance and correction corpora are part of long-term linguistic accountability. In AI procurement, institutions should clarify before contracting their rights to evaluation, incident notification, notice of model changes, audit access, data deletion, exit and portability, and allocation of liability. Procurement itself is a stage of governance. In governance of model updates, significant behavioural changes should trigger revalidation. Even under the same product name, weights, retrieval systems and safety layers may change. Approval should therefore be version-specific. Incident reporting should include near misses. If harm did not occur because an operator caught the problem in time, the event is still evidence of future risk. A safety culture should not count only realised accidents. A public-sector transparency register can publish the purpose, responsible agency, vendor, impact category, review date, complaint path and audit summary for deployed AI systems, with privacy and security exceptions kept narrow. A simple rule for individual users is: on an important claim, “show me the source” is not enough—open the source, check the quotation, verify the date and version, seek an alternative source, and consult an expert when necessary. Treat AI as a navigator, not as the final witness. Chapter Conclusion In the AI age, evidence no longer means merely “where is the link?” It includes the entire chain by which a claim was produced: which sources, model and version were used, what human decisions were involved, what uncertainty remained, and how re- examination is possible. This chain is the modern form of evidential accountability. Accountability likewise is not merely blame. Without role-specific duties, competent oversight, traceability, monitoring, incident response, correction and appeal, phrases such as “human in the loop” or “ethical AI” may remain empty. One of the most dangerous sentences may be: “the system said so.” It occupies the place of a reason without being a reason, and it assumes the form of authority without legitimacy. The proper questions are: why was the system configured this way, what is the evidence, who reviewed it, who had decision authority, and who will correct the harm if harm occurs? Indian theories of pramāṇa and Mithila’s Nyāya–Navya-Nyāya tradition offer an important lesson for modern AI governance: subject knowledge claims to examination of instrument, defect, doubt, pervasion, hidden conditions, opposition and defeat. Modern audit is not a mechanical copy of this older discipline, but a parallel reconstruction. The maxim of this chapter is: “Do not treat an AI answer as evidence; construct the provenance chain of evidence. Do not let a machine decision dissolve responsibility; construct a chain of roles, review, correction and appeal.” गजेन्द्र ठाकु रक समानान्तर दर्शन — खण्ड २ PART XIV — PARALLEL PHILOSOPHY IN THE TWENTY-FIRST CENTURY