arXiv
Large language models (LLMs) may possess internal signals correlated with answer correctness, yet these signals are not necessarily expressed faithfully through verbalized confidence. We argue that LLM metacognition should therefore be viewed not only as an intrinsic model capability but also as an elicited behavior shaped by external rules, incentives, and interaction contexts. Motivated by this perspective, we formulate confidence elicitation as a game theoretic mechanism design problem. Building on classical proper scoring principles, we introduce a correctness enhanced Brier payoff as a training free protocol that explicitly communicates the consequences of answer correctness and confidence misreporting. Its normative optimum corresponds to answer specific truthful probability reporting for an ideal expected payoff maximizer. We further introduce Batch contrastive elicitation, which presents multiple questions jointly and provides a local reference set to compare knowledge familiarity, ambiguity, reasoning complexity, and potential error sources. The payoff specification determines what confidence behavior should be rewarded, while batch contrast enriches the context from which confidence is formed. We evaluate these mechanisms on four benchmarks and five open-weight and proprietary LLMs. The results show that explicit scoring incentives and batch contrast improve confidence calibration, correct-versus-incorrect discrimination, and high confidence reliability, with more consistent benefits for stronger models and challenging factual tasks. These findings demonstrate that appropriately designed external mechanisms can help transform latent uncertainty signals into more reliable metacognitive behavior without parameter updates or access to model internals.