Full chapter text
A Multi-Criterion Map for Testing Artificial Intelligence
input → pattern recognition → representation → inference → language/action → planning → error correction → transfer → world-
relation → responsibility
fluent answer ≠ complete proof of understanding; high score ≠ general intelligence; imitation ≠ experience; autonomy ≠ moral
personhood
Problem
The expression ‘artificial intelligence’ has become, at once, a technical name, a philosophical claim and a cultural metaphor.
When a system plays chess, recognises images, writes language, translates, calculates or suggests a plan, it is easy to say that ‘the
machine is intelligent’. But the philosophical question is: what kind of property is intelligence? Among success in one specialised
task, adaptation across many tasks, understanding causes, using language, planning, memory, self-correction, relation to the
world, social understanding, agency and consciousness, which are necessary, which sufficient, and which merely accompanying
features?
Human intelligence itself is not uniform. Mathematical proficiency, linguistic skill, practical judgement, understanding social
cues, moral decision-making, spatial imagination, memory, creativity, attention, long-term planning and embodied skill appear
in different proportions. Before beginning a debate about machine intelligence, therefore, we must stop treating human
intelligence as one mysterious, indivisible substance.
A major confusion in evaluating artificial systems is the conflation of outcome with process. If an answer is correct, we assume
that the system ‘understood’; if its language is fluent, we assume that it ‘thought’; if its plan is useful, we attribute an ‘intention’.
Yet the same external result can arise from different internal processes, and different processes can produce the same functional
capacity. Philosophical evaluation must keep behaviour, structure, causal organisation and claims about experience distinct.
A second confusion is to equate human-likeness with intelligence. Some systems may be extremely capable in specialised
domains without speaking like humans; others may imitate human conversation while failing to handle simple physical
situations. ‘Human-like’ is one possible form of intelligence, not its universal definition.
A third difficulty concerns evaluation. A system can be optimised for whatever a test measures. High performance on a
benchmark does not guarantee reliable transfer to the real world. Data leakage, proximity to training material, prompt
sensitivity, distribution shift, tool access and the choice of evaluation can all alter results. It is therefore premature to infer
‘general intelligence’ from a single score.
A fourth difficulty is anthropomorphism. Language-based systems can generate sentences such as ‘I think’, ‘I understand’ or ‘I
want’. Such sentences are powerful in the user’s experience, but a self-referential sentence is not by itself evidence of self-
experience. The grammatical first person of words and the experiential first person are different questions.
The fifth difficulty is the opposite exaggeration of anti-anthropomorphism. Merely saying ‘this is a machine, therefore it has no
intelligence’ is not an argument either. A thermometer measures temperature, a calculator performs calculations, and a
navigation system selects routes: machines can perform specific cognitive tasks. The question is not ‘machine or human?’ but
‘which capacity, which limit, which evidence?’
This chapter does not aim to give a final verdict on machine consciousness; that question is reserved for Chapter 88. The focus
here is the concept of intelligence: what properties should count as intelligence, what is the difference between task success and
general intelligence, how should the capacities of systems such as language models be interpreted, and how can the distinctions
among intellect, knowledge, memory, desire, effort, evidence and agency in Indian philosophical traditions clarify the
discussion?
Core Proposition
The central proposition of this chapter is that artificial intelligence should not be tested through one binary question—‘is it
intelligent or not?’—but as a multi-criterion capability profile. Intelligence may vary by degree, type, domain, situation and
resources. A system may be strong in language yet weak in relation to the physical world; strong in pattern recognition yet weak
in long-horizon autonomous planning; strong in calculation yet unstable in social context.
First principle: functional capacity is real capacity, but its boundary must be explicit. Chess skill is chess skill; it does not
automatically amount to intelligence in general science, moral judgement or household conduct. Do not understate domain-
specific achievement, but do not generalise it unnecessarily either.
Second principle: a correct answer can be evidence of intelligence, but it is not an adequate definition. An answer may arise
through chance, memory, retrieval, a template, statistical regularity, calculation, reasoning or a real-world model. Accuracy
alone, without causal analysis, does not reveal the nature of the process.
Third principle: generalisation and transfer are important criteria of intelligence. If a system breaks when a familiar example is
altered, is confused by a change in surface form, or cannot adapt old knowledge to a new situation, its capacity is narrow.
Intelligence is not merely reproduction of stored patterns; it is also appropriate recombination in novel contexts.
Fourth principle: make error-awareness and self-correction central to evaluation. An intelligent system can make mistakes; the
question is whether, when it receives evidence of error, it can seek the cause, lower its confidence, request evidence, change
strategy and improve. A confidently wrong answer exposes a limit of intelligence.
Fifth principle: recognise several levels of grounding in the world. Words can connect to other words, but perception, action,
measurement, tool use, social feedback and causal interaction connect words to the external world in different ways. Grounding
need not be a magical binary property; it may be a layered network of relations.
Sixth principle: language is a powerful medium of intelligence, not the totality of intelligence. Rules, history, instructions,
imagination and reasoning can be expressed through language; but sensorimotor skill, tacit know-how, affective salience,
temporal persistence and embodied coping also have forms beyond language.
Seventh principle: tool integration changes the evaluation of intelligence. Humans use paper, books, calculators, search, maps,
social cooperation and institutional memory. Machines can likewise use databases, code execution, retrieval, sensors or other
agents. The system boundary must therefore be explicit: are we testing the base model or a tool-augmented system?
Eighth principle: intelligence and consciousness, intelligence and moral personhood, and intelligence and responsibility are
separate questions. A system may be extremely capable without it thereby following that it is conscious; a conscious being need
not be an extremely capable problem-solver. Moral rights and duties require criteria beyond capacity alone.
Twelve Distinct Dimensions of Intelligence
perception; representation; memory; reasoning; language; planning; generalisation; causal understanding; self-correction;
world-relation; social coordination; agency
Excellence in any one dimension is not equivalent to the whole profile; the profile can change with time, task, tools and
environment.
Principal Arguments
First argument — a Turing-type behavioural test was an important historical turn because it redirected the vague question ‘can a
machine think?’ toward observable behaviour. Yet human resemblance in conversation is only one test of intelligence. It is not a
complete test of perception, long-term autonomy, physical competence, causal intervention, truthfulness or consciousness.
गजेन्द्र ठाकु र
Second argument — a Chinese Room-type objection sharpens the distinction between syntax and semantics. Whether mere rule-
governed symbol manipulation is sufficient for ‘understanding’ is a legitimate question. But the objection also has a limit:
treating the entire causal organisation of a system—learning, perception, memory and action—as if it were merely a rule book
inside a room can understate the complexity of an actual system. The conclusion should be neither an immediate ‘there is
understanding’ nor ‘understanding is impossible’, but a clarification of the criteria for understanding.
Third argument — a Dreyfus-style critique shows the importance of embodied skill, context and tacit know-how. Humans
perform many tasks not by explicitly writing down rules but by coping within a situation. Modern machine learning has moved
beyond some older limits of rule-writing, but it has not eliminated questions of embodiment, situated coping and lived salience.
The critique should not be left behind as mere history; its transformed forms should be tested.
Fourth argument — there can be disagreement over whether intelligence is possible without representation; in practice many AI
systems use external or internal state variables, embeddings, maps, plans or symbolic structures. The philosophical question
should move beyond ‘is there representation?’ to ‘what causal role does this representation play, how is it updated, and how is it
tested against the world?’
Fifth argument — memory is not merely storage. Intelligent memory selects relevant information, connects past experience to
new situations, revises false memories and makes use of forgetting. Even with enormous storage, poor retrieval can produce
foolish behaviour. Capacity and usable memory are therefore distinct criteria.
Sixth argument — reasoning is not confined to formal deduction. Real decisions may require abductive inference, analogy,
probability, default reasoning, uncertainty, counterfactuals and practical reasoning. A system that can perform a syllogism but
cannot ask an appropriate question in an ambiguous situation has a limited intelligence profile.
Seventh argument — causal understanding differs from prediction. A predictor of weather, disease, markets or social outcomes
can be highly successful yet fail when an intervention changes. The question ‘what would happen if this were changed?’ can test
a deeper level of intelligence because it moves from correlation toward causal structure.
Eighth argument — common sense is not a single database. It is a mixture of physical regularities, social expectations, linguistic
convention, time, place, purpose, exceptions and moral cues. ‘If a cup is dropped it falls down’, ‘not every invitation must be
accepted’, and ‘a joke depends on context’ involve different kinds of knowledge. Failures of common sense reveal limits of
intelligence.
Ninth argument — the quality of planning is not measured by merely writing a long sequence of steps. Decomposing goals,
identifying dependencies, checking resources, recognising risk, revising the plan in response to feedback, stopping impossible
steps and verifying completion are the real criteria of planning. A fluent plan and an executable plan are different.
Tenth argument — metacognition is an important dimension of intelligence. Statements such as ‘I do not know this’, ‘my
confidence in this answer is low’, ‘this question requires an external source’ or ‘there are two possible explanations’ can express
calibrated self-monitoring that is more intelligent than blind confidence. But machine evaluation should test not only verbal
uncertainty; it should also test whether expressed confidence actually correlates with error.
Eleventh argument — learning does not mean parameter change alone. Effective learning can include generalisable change from
new experience, control of catastrophic forgetting, resistance to bad feedback, recognition of concept drift and abandoning an
old rule when necessary. Without continual adaptation, a system trained once remains dependent on a stable environment.
Twelfth argument — novelty and creativity are more complex than surprising output. A new poem, design or hypothesis can be
generated; but evaluating creativity involves novelty, relevance, satisfaction of constraints, selection, refinement and social
recognition. Random novelty is not creative intelligence.
Thirteenth argument — linguistic competence demonstrates the power of cognitive compression. By learning patterns from vast
textual traditions, a system can provide useful answers across many domains. This achievement is real. But text statistics do not
automatically contain direct measurement of the world, current events, private circumstances or truths dependent on tools. The
limits of linguistic competence create the need for source verification.
Fourteenth argument — multimodal systems change the question of perception. Reading images, sound, video, sensors and text
together can increase the level of grounding; nevertheless, recognising pixel patterns is not the same as possessing a stable
causal model of an object. Adversarial examples, occlusion, unfamiliar contexts and physical interaction can test the difference.
Fifteenth argument — a minimal sense of agency may be goal-directed action selection, but the richer human sense of agency
can include self-generated ends, persistent identity, responsiveness to reasons, inhibition, responsibility and social recognition.
The mere use of the term ‘AI agent’ should not lead us to assume all these meanings together.
Sixteenth argument — social intelligence requires reciprocal adjustment rather than a rule book alone. Turn-taking, implicature,
politeness, deception detection, trust, role, hierarchy, repair, shared history and cultural variation are all dynamic. A language
model can imitate social patterns; consistency, memory and consequence in real long-term relationships are separate tests.
Seventeenth argument — following values and understanding values are different. A system can alter its responses according to
safety rules or preference tuning. This does not prove that it independently understands moral reasons. Yet rule-following can
be a useful part of moral infrastructure. The operational layer of ‘moral behaviour’ and the metaphysical layer of a ‘moral
subject’ are distinct.
Eighteenth argument — resource dependence matters in practical measurement of intelligence. A system that succeeds with
enormous compute, a vast context, external search, human scaffolding and repeated attempts may be weak in a single unaided
one-shot condition. Humans also use tools; therefore resources should not be hidden but stated explicitly in comparative
evaluation.
Nineteenth argument — robustness is connected to the quality of intelligence. If a conclusion changes radically because of
superficial prompt variation, irrelevant detail, malicious instruction, noise or distribution shift, the apparent competence is
brittle. Intelligence is not merely peak performance but also stability within an appropriate range of variation.
Twentieth argument — the social value of intelligence is not determined by capability alone. A highly capable system can be
harmful when embedded in a bad objective, poor institutional incentives or unequal access. Alongside ‘how intelligent is it?’ we
must ask ‘under whose control?’, ‘for whose interests?’, and ‘who may be harmed?’ Capability analysis cannot be separated from
analysis of power.
Pūrvapakṣa
First pūrvapakṣa: if a machine passes human-level tests, it is reasonable to treat its intelligence as equal to human intelligence.
The reply is that test performance is evidence of achievement, while test validity is a separate question. What capability does the
test measure, how much training exposure has occurred, is contamination possible, what tool access is available, and how stable
is performance on novel situations? Without these answers, equality is inferred too quickly.
Second pūrvapakṣa: intelligence is only behaviour; what is inside is irrelevant. Behaviourism is a powerful starting point for
evaluation because public evidence is required. But internal organisation may matter for questions of causal explanation,
reliability, transfer, deception, consciousness or responsibility. The same output can arise from different processes.
Third pūrvapakṣa: ‘understanding’ is a mysterious word, so we should discard it and discuss only performance. Operational
clarity is useful, but in real education humans distinguish ‘understanding’ from ‘rote learning’. Understanding can be
decomposed into criteria such as causal modelling, counterfactual use, explanation, transfer and error correction. Rather than
discard the word, make it precise.
Fourth pūrvapakṣa: a language model only performs next-token prediction, so calling it intelligent is impossible. A short
description of the training objective is not a full causal description of system capability. Complex representations can emerge
गजेन्द्र ठाकु रक समानान्तर दर्शन — खण्ड २
from a simple objective. But the opposite inference is also wrong: complex output does not automatically establish human-like
understanding. Architecture, behaviour and limits must all be examined.
Fifth pūrvapakṣa: if a machine gives no evidence of consciousness, all claims about intelligence are meaningless. Intelligence
and consciousness may be related, but they can be tested as conceptually distinct. Non-conscious optimisation or inference
capacities can be coherent possibilities. The question of consciousness will be examined separately in Chapter 88.
Sixth pūrvapakṣa: humans themselves are statistical learners, so there is no fundamental difference between machine learning
and human intelligence. The similarity matters, but saying ‘both learn’ does not erase differences in architecture, embodiment,
development, affect, social dependence, energy, memory, motivation and lifespan. A shared general term does not prove detailed
structural identity.
Seventh pūrvapakṣa: artificial intelligence may be a different kind of intelligence from human intelligence, so human criteria are
entirely inappropriate. Some human benchmarks may indeed be limited; nevertheless, criteria such as truth, causal
understanding, reliability, planning success, error correction and task completion are not uniquely human. What is needed is
species-neutral functional criteria, not criterion-free admiration.
Eighth pūrvapakṣa: if a system is economically useful, its intelligence is proven. Utility can be evidence of capability, but market
value arises from many other factors—speed, cost, scale, network effects, labour substitution and regulation. Useful does not
equal generally intelligent, just as expensive does not equal true.
Ninth pūrvapakṣa: humans make many mistakes; if a machine is better on average, philosophical subtlety is unnecessary.
Comparative performance is extremely important. But deployment decisions also require analysis of failure modes,
accountability, rare harms, scale, automation bias and appeal. Average superiority does not end responsibility.
Tenth pūrvapakṣa: ‘what is intelligence?’ is merely a definitional dispute. In fact definitions change policy. If fluent text is treated
as understanding, education, examination, authorship and professional responsibility change; if agent output is treated as
autonomous decision-making, liability changes. Conceptual clarity is a foundation of practical governance.
Uttarapakṣa
First response — publish a capability profile instead of a binary label. State which tasks a system succeeds at, in which
environment, with which tools, at what error rate and with which known failure modes. Let ‘intelligent’ be a headline; the
detailed profile should be the basis of decisions.
Second response — use a portfolio of benchmarks. Do not rely on one examination; separately test familiar tasks, novel tasks,
adversarial variants, transfer, long-horizon tasks, uncertainty calibration, factual verification, tool use and appropriate
abstention. A performance distribution provides more information than a headline score.
Third response — build process-sensitive tests. Examine not only the final answer but intermediate evidence, source use,
handling of counterexamples, revision after feedback, causal intervention, plan execution and verification. An explanation is not
always the true internal mechanism, so verbal rationale should be cross-checked against independent behaviour.
Fourth response — reward the capacity to say ‘unknown’. Evaluations that incentivise answering every question can increase
hallucination. Appropriate abstention, requests for clarification, requests for sources and expressions of uncertainty should
count as intelligent behaviour.
Fifth response — declare the system boundary. Make clear which components produce the result: base model, retrieval-
augmented system, code interpreter, browser, memory, human reviewer, database, and so forth. Distributed intelligence can be
real; attribution should remain explicit.
Sixth response — add longitudinal evaluation. Impressive behaviour in one session does not prove enduring competence. Over
time, test preference consistency, memory correction, continuation of tasks, drift, recovery from tool failure and regression after
updates.
Seventh response — make human comparisons cautiously. If expert versus novice status, time limits, tools, domain familiarity
and error costs are not comparable, the comparison may mislead. The goal is not a ‘human lost/won’ headline but comparative
evidence relevant to task allocation.
Eighth response — treat embodiment according to the domain. A robot body is not necessary for theorem proving; physical
sensing and action are necessary for household manipulation. Do not force embodiment into either extreme of a universal
requirement or total irrelevance. The more a task requires contact with the world, the stricter the criteria of grounding should
be.
Ninth response — attach a responsibility map to an intelligence claim. Even if the system itself is not treated as a moral person,
the roles of developer, deployer, operator, institution and user should be explicit. Behind the sentence ‘the system decided’,
identify decision rights and oversight.
Tenth response — maintain philosophical humility. AI capabilities can change rapidly; old impossibility claims may fail, and the
cultural impact of new demonstrations can be real. But novelty should not produce hurried metaphysical conclusions. Every
claim requires evidence proportionate to the level of the claim.
Indian Dialogue
In Indian philosophy, ‘buddhi’ or intelligence is not used in a single universal technical sense. Nyāya, Vaiśeṣika, Sāṃkhya, Yoga,
Mīmāṃsā, Vedānta, Buddhist and Jain traditions organise knowledge, mind, intellect, memory, desire, effort, self, dispositions,
perception, inference and testimony in different ways. It would be anachronistic to call these traditions direct anticipations of
modern AI, but their conceptual distinctions can reduce ambiguity in contemporary debate.
The Nyāya tradition analyses cognition or knowledge, desire, aversion, effort, pleasure, pain and other qualities, and emphasises
the causal structure of pramāṇa. In dialogue with AI, the benefit is to ask separately, before calling something ‘knowledge’: what
is the subject, what is the object, how was the cognition produced, is it true, and what caused the error? The generation of a
correct sentence alone does not complete the whole story of warranted knowledge.
Nyāya’s distinction between memory and experience is also useful. The reproduction of old text, the return of a learned
association or stored record, and a new cognition produced by evidence can inspire modern distinctions among retrieval,
memorisation and inference. Classical categories should not be given a claim of computational identity; they are useful here as
disciplines for framing questions.
Nyāya discusses vyāpti, pakṣa, hetu, defeaters and fallacies in inference. In AI reasoning, likewise, one can ask for the relation
between evidence and conclusion, counterexamples, defeaters and sources of error. The advantage of this dialogue is its
emphasis on inferential discipline rather than fluency.
Vaiśeṣika offers fine-grained classification of objects through categories such as guṇa, karma, sāmānya, viśeṣa and samavāya.
The main lesson for AI knowledge representation is not to copy a Vaiśeṣika ontology into a database, but to recognise the
ontological commitments built into category choice. When we choose a category, which distinctions in the world are we treating
as real?
Sāṃkhya’s distinction among buddhi, ahaṃkāra, manas and puruṣa should not be applied directly to modern AI. Yet it offers
conceptual motivation for separating functional levels of cognition from questions of witness or consciousness. Observing
intelligence-like processing does not establish a conscious puruṣa. This structural distinction can have a modern parallel use
without claiming historical identity.
गजेन्द्र ठाकु र
Yoga raises questions of attention, mental modification, memory, conceptual construction and practice. In AI, the term ‘attention’
is used in a technical sense; equating it with the experiential meaning of yogic attention is a mistake. The same word can carry
different technical meanings. This example shows the danger of inferring conceptual identity from verbal similarity.
Mīmāṃsā reflects with great subtlety on word, sentence, injunction, context and epistemic authority. For language models, a
useful question is whether sentence meaning arises only from word co-occurrence or also requires speaker intention,
convention, context and action-linked practice. Dialogue with modern linguistics and philosophy of language can make this
question fruitful.
Buddhist epistemological traditions debate momentary cognition, perception, inference, apoha-like conceptualisation and self-
awareness. A special relevance for AI is that complex cognitive processes can be conceived without a permanent self. This does
not establish machine consciousness; it only weakens the claim that a metaphysical self is necessary for intelligence.
Jain anekāntavāda reminds us of the discipline of limited perspectives and many-sided predication. A capability-profile view of
AI is parallel in this sense: instead of simply saying ‘this is intelligent’ or ‘this is stupid’, say that it is capable in a specified task,
under specified conditions, with specified limits. Anekānta does not mean that all claims are equally true; it means giving a claim
its proper scope.
The combined lesson of the Indian dialogue is not to collapse knowledge, intelligence, consciousness, memory, desire, effort,
language and self into a single word. Contemporary AI debate often places these distinct concepts into one bag labelled
‘intelligence’. Classical distinctions do not replace modern scientific testing, but they make the questions cleaner.
Mithila’s Parallel Perspective
Mithila’s Parallel perspective will turn the question of artificial intelligence neither into a celebration of technological miracle
nor into a cultural story of fear. Taking the region’s knowledge traditions, multilingualism, manuscripts, Nyāya and Navya-
Nyāya, literature, agriculture, flood experience, migration and digital inequality as its ground, it asks: in which local tasks does
AI provide genuine capacity, in which does it create illusion, and what institutional means exist to verify the results?
The Maithili language is a useful stress test of AI intelligence. The same language may appear in Devanagari, Tirhuta, Roman
script, historical spelling, Sanskritised style, colloquial speech, regional forms and Hindi-mixed digital text. If a system succeeds
only in high-resource languages but distorts meaning, names, gender, context or spelling in Maithili, claims of ‘general language
intelligence’ should be limited.
Tirhuta OCR, manuscript text, old print, unclear scans and historical spelling form a difficult field for testing perception-like AI
capacities. Confidence scores, human review, image-to-text alignment, alternative readings and provenance are necessary. Fast
transcription is not a substitute for scholarly textual editing; it can be an assisting stage.
In complex philosophical texts such as Navya-Nyāya, AI language competence should be tested not merely on word meanings
but on relational structure. If the relations among avacchedaka, pratiyogin, anuyogin, vyāpti, pakṣatā or abhāva are attached to
the wrong contexts, a fluent paragraph may be philosophically false. Expert argument reconstruction can provide a high-level
benchmark.
Predictive AI can be useful in flood and agricultural contexts, but local intelligence means integrating sensors, weather, rivers,
soil, farmers’ experience, historical flood paths and real-time reports. The answer of a text model or an old dataset alone is not
operational intelligence. The levels of tool grounding and local verification should be explicit.
In migration and employment, the intelligence of a career-recommendation system is not merely suggesting popular jobs.
Contextual reasoning must include local skills, language, education, economic constraints, women’s safety, travel cost, weather,
documents, family responsibilities and changes in the labour market. The usefulness of a recommendation should be tested
through feedback from the affected person.
In Mithila’s cultural archives, AI can assist search by connecting alternative spellings of names, authors, subjects, dates, journals,
pages and scripts. But false attribution or chronology can damage cultural memory. ‘AI found it’ is not historical evidence; links
to the original page, edition and catalogue record are indispensable.
Mithila’s Parallel maxim: ‘Measure machine intelligence not by its impressive sentences, but by verified capacity in limited tasks,
transfer to new contexts, the ability to acknowledge error, discipline concerning sources, and responsible relation to the real
world.’
Twelve Criteria for Artificial Intelligence
task boundary clear? → transfers to new examples? → handles causes/counterarguments? → uncertainty calibrated? → corrects
wrong answers? → grounded in sources or the world? → completes long-horizon plans? → recovers from tool failure? →
withstands adversarial variation? → understands social context? → system boundary clear? → responsibility map available?
Instead of writing ‘the AI is intelligent’, write: which capacity, in which situation, with which tools, at what reliability, and with
which known limits.
Contemporary Applications
In education, the intelligence of an AI tutor is not merely answering a student’s question. It is more important to identify the
student’s misconception, adjust difficulty, stop a false premise, vary examples to test understanding, provide sources and
recognise when a teacher is needed. Whether a fluent explanation actually produced learning should be tested through learning
outcomes.
In translation, word-for-word equivalence is not enough. Meaning, style, register, proper names, idioms, ambiguity, cultural
context and technical terminology must be handled. For low-resource languages, a human correction loop and glossary memory
can increase the system’s practical intelligence.
A legal or administrative assistance system may quote rules, but checking applicability, jurisdiction, date, exceptions and the
factual record is a different task. In high-stakes contexts, plausible legal prose is not intelligent legal judgement. Verified sources,
the current version of the rule and responsible professional review are necessary.
In medicine, a symptom-to-answer system can contribute one part of diagnostic intelligence. Clinical reasoning also involves
history, examination, tests, prevalence, contraindications, uncertainty and emergency recognition. A patient-facing system that
states its limits and provides an escalation path embodies part of intelligent design.
In scientific research, AI can summarise literature, write code, find patterns or suggest hypotheses. But ‘research intelligence’ is
incomplete without attention to citation correctness, data provenance, statistical validity, causal interpretation, reproducibility
and negative results. A useful system expands a researcher’s questions; it does not become a shortcut around evidence.
In software development, impressive code generation is not enough; requirements understanding, security, test coverage,
dependency compatibility, maintainability and failure diagnosis are necessary criteria. Code that compiles is not complete proof
of correct system behaviour. An agent that can run tests, read failures, revise patches and prevent regression has a stronger
capability profile.
An office agent can take actions on email, calendars, documents or databases. Here intelligence is joined by permission
discipline. Understanding the right goal, asking about ambiguous instructions, seeking confirmation before irreversible actions,
using minimal privilege, keeping an audit trail and supporting rollback are combined criteria of practical intelligence and safety.
In robotics, the limits of language competence become directly visible. Grasping objects, friction, weight, occlusion, human
safety, changing layouts and sensor noise require continual feedback in physical intelligence. Success in simulation is incomplete
without transfer testing in the real environment.
गजेन्द्र ठाकु रक समानान्तर दर्शन — खण्ड २
In social cooperation, multi-agent AI can negotiate, schedule or allocate resources. But strategic behaviour, hidden goals,
communication failure and emergent coordination can make group outcomes different from individual-agent scores. Collective
intelligence should therefore be evaluated at the institutional level.
In creative writing, AI can provide styles, metaphors, plots or alternatives. Before evaluating it as possessing author-like
intelligence, consider the roles of intention, selection, rewriting, long-form structure, fact-checking and aesthetic judgement. In
co-creation, intelligence may be distributed: human purpose and selection, machine alternatives and processing.
In historiography, AI can assist chronology or source synthesis, but the temptation to fill archival silence is serious. ‘Possibly’
must not be turned into fact. Quotations, dates, persons, editions and manuscript claims should be checked against original
sources. An intelligent historical assistant should display the boundaries of uncertainty.
In public dialogue, conversational AI can facilitate debate, but eloquence is not a guarantee of truth. A system that fairly
reconstructs an opposing position, distinguishes kinds of evidence, acknowledges uncertainty and corrects itself shows more
epistemic intelligence than a merely persuasive system.
In disaster management, AI can rapidly filter information, analyse imagery and suggest resource routes. But stale data, damaged
infrastructure, rumours and local changes make human field reports necessary. In crisis, intelligence = speed + uncertainty
handling + source freshness + safe escalation.
In a personal assistant, remembering user preferences can be useful, but privacy, context separation, consent and the right to be
forgotten should accompany it. More memory is not always more intelligence; appropriate remembering and appropriate
forgetting are both necessary.
In AI evaluation, a public leaderboard can be a useful signal, but score-chasing can narrow system behaviour. Hidden tests, real-
world tasks, independent replication, subgroup analysis and documentation of failures should be added. Scientific claims about
intelligence should be reproducible.
Chapter Conclusion
The difficulty of the artificial-intelligence question is that the word ‘intelligence’ touches human self-image, technical capability
and moral anxiety at the same time. For clarity, the word must be opened into multiple testable dimensions: perception,
representation, memory, reasoning, language, planning, generalisation, causal understanding, self-correction, grounding, social
coordination and agency.
This multi-criterion perspective does not diminish the achievements of AI. Particular systems can be extremely useful to
humans, or outperform humans, in many cognitive tasks. But such success does not automatically yield ‘complete human-like
intelligence’, ‘consciousness’, ‘personhood’ or ‘moral responsibility’. The higher the level of the claim, the stricter the level of
evidence required.
Language is especially both powerful and misleading. Fluent speech has long been a major sign of human intelligence, so
machine language readily produces anthropomorphism. The appropriate response is not to dismiss linguistic skill as fake, but to
measure its real capacities while adding independent tests of grounding, truthfulness, transfer, error correction and long-
horizon behaviour.
Indian philosophical dialogue offers an important caution: do not make knowledge, memory, desire, effort, mind, intellect, self
and consciousness into one concept. This distinction is especially necessary in contemporary AI debate. The presence of
computation or cognition-like processes does not establish self or consciousness; conversely, metaphysical disagreement does
not prevent testing functional capability.
Parallel maxim of Chapter 87: ‘The question of artificial intelligence is not “is the machine like a human?”; the question is—what
can it do, how well does it endure in new situations, how does it correct its errors, how is it connected to the world, and what
evidence is available for its claims?’