← Lectures·L13

Philosophy of AI: Ethical & Epistemological Issues

Robertson · 1 concepts · 6 questions

Concept 1 / 1

Key Concepts to Memorize

3 types of DNN opacity:


1. Intentional secrecy — deliberate non-disclosure by companies (Burrell 2016)

2. Technical illiteracy — users/stakeholders lack the expertise to understand (Burrell 2016)

3. Algorithmic-level opacity (ALO) — the operations of DNNs are inscrutable even to their own developers due to sheer complexity


Algorithmic-level opacity (most important):


  • Developers can inspect weights and bias parameters
  • But they CANNOT explain HOW the model maps inputs to outputs
  • They do not grasp "which higher-level mathematical structures and processes these parameters implement" (Zednik & Boelson 2022)

xAI (Explainable AI) — the controversy:


  • Methods like LIME (surrogate models, saliency maps) are proposed as solutions to opacity
  • Key criticism: xAI provides post-hoc explanations — they may NOT faithfully represent the original black box model
  • Counter-argument: human experts also give post-hoc explanations of intuitive judgements (Zerilli et al. 2019) → "biological double standard"

Responsibility gaps:


  • "The highly autonomous behavior of AI systems, for which neither the programmer, the manufacturer, nor the operator seems to be responsible" (Königs 2022)
  • AI decisions → unclear who is liable → "problem of many hands"

Instrumental Convergence Thesis:


  • Agents (artificial or otherwise) will seek to fulfil instrumental goals (resource acquisition, power-seeking) to fulfil larger-scale goals
  • Even if an AI's top goal is benign, it may conflict with human interests through instrumental goals
  • Classic example: Bostrom's paperclip maximizer — a superintelligence manufacturing paperclips could convert all matter on Earth into paperclip manufacturing facilities

Trustworthy AI:


  • EU AI Act and UNESCO advocate for "trustworthy AI"
  • Philosophical question: are AI systems appropriate objects of trust? Or just objects of reliability?
  • Trust (unlike reliability) can be betrayed — can AI systems enter such an interpersonal relation?

LLMs as cognitive agents? Two camps:


  • Observational support (Cappelen & Dever 2025): We use intentional terms (believes, knows, acts) to describe LLMs → evidence they have minds
  • Skeptical: LLMs hallucinate in ways incompatible with genuine understanding (unlike human mistakes)

MIT AI Lab context:


  • Co-founded in 1959 by Marvin Minsky and John McCarthy (originator of the term "artificial intelligence")