Philosophy of AI: Ethical & Epistemological Issues
Robertson · 1 concepts · 6 questions
Concept 1 / 1
Key Concepts to Memorize
3 types of DNN opacity:
1. Intentional secrecy — deliberate non-disclosure by companies (Burrell 2016)
2. Technical illiteracy — users/stakeholders lack the expertise to understand (Burrell 2016)
3. Algorithmic-level opacity (ALO) — the operations of DNNs are inscrutable even to their own developers due to sheer complexity
Algorithmic-level opacity (most important):
- Developers can inspect weights and bias parameters
- But they CANNOT explain HOW the model maps inputs to outputs
- They do not grasp "which higher-level mathematical structures and processes these parameters implement" (Zednik & Boelson 2022)
xAI (Explainable AI) — the controversy:
- Methods like LIME (surrogate models, saliency maps) are proposed as solutions to opacity
- Key criticism: xAI provides post-hoc explanations — they may NOT faithfully represent the original black box model
- Counter-argument: human experts also give post-hoc explanations of intuitive judgements (Zerilli et al. 2019) → "biological double standard"
Responsibility gaps:
- "The highly autonomous behavior of AI systems, for which neither the programmer, the manufacturer, nor the operator seems to be responsible" (Königs 2022)
- AI decisions → unclear who is liable → "problem of many hands"
Instrumental Convergence Thesis:
- Agents (artificial or otherwise) will seek to fulfil instrumental goals (resource acquisition, power-seeking) to fulfil larger-scale goals
- Even if an AI's top goal is benign, it may conflict with human interests through instrumental goals
- Classic example: Bostrom's paperclip maximizer — a superintelligence manufacturing paperclips could convert all matter on Earth into paperclip manufacturing facilities
Trustworthy AI:
- EU AI Act and UNESCO advocate for "trustworthy AI"
- Philosophical question: are AI systems appropriate objects of trust? Or just objects of reliability?
- Trust (unlike reliability) can be betrayed — can AI systems enter such an interpersonal relation?
LLMs as cognitive agents? Two camps:
- Observational support (Cappelen & Dever 2025): We use intentional terms (believes, knows, acts) to describe LLMs → evidence they have minds
- Skeptical: LLMs hallucinate in ways incompatible with genuine understanding (unlike human mistakes)
MIT AI Lab context:
- Co-founded in 1959 by Marvin Minsky and John McCarthy (originator of the term "artificial intelligence")