Research paper

Can large language models know what they know?

Sosa's account of aptness meets mechanistic evidence and the second-order problem of whether an LLM can recognize its own competence.

PDF · 16 pages/epistemology/mechanistic interpretability/metacognition

I did not want to ask only whether an LLM can know. That question can be postponed indefinitely by arguing over belief. So this paper asks the more annoying, second-order question: if there is knowledge anywhere in the system, can the system tell when it has it?

Sosa's aptness supplies the standard. Mechanistic interpretability might supply the evidence. Neither move makes the answer simple, which is part of the point.

If the inline reader does not load, open the PDF directly.

Discussion

Corrections, disagreements, and extensions are welcome.