Short paper
Sometimes, LLMs think
A reliable-faculties route to machine thought, and a role for mechanistic interpretability in separating competence from coincidence.
Some arguments about machine thought begin by deciding what thought must feel like from the inside. This one starts lower down, with reliable faculties and mechanism.
It asks whether parts of a language model can succeed for the right internal reasons — and whether interpretability can tell the difference between competence and coincidence.
The paper
Download PDF ↓If the inline reader does not load, open the PDF directly.
Discussion
Corrections, disagreements, and extensions are welcome.