A connected research program · Shiyu Ni · 2024–2026
Knowledge boundaries. Honest AI.
Knowing an answer is not the same as knowing when to trust it.
Our work connects knowledge-boundary evaluation, the explanation of overconfidence, and honesty alignment—with one goal: helping models make better decisions about their own answers.
Recognizing its limits helps a model decide when to answer, seek evidence, or defer. We study how to make this judgment reliable—not merely how to attach a confidence score.
A conceptual overview of the research direction; plots are schematic. Click to read at full size.Open original ↗
02 / Developing the argument
From observing a gap to learning reliable self-assessment.
Across this program, one question persists: how can a model’s judgment of its own answers better reflect what it can actually do?
Chapter 01
Understanding knowledge boundaries
We began with a practical question: when should a model rely on its own knowledge, and when should it retrieve? Answering it requires measuring the gap between what a model knows and what it thinks it knows.
Our ACL 2024 Findings study makes this gap measurable and identifies overconfidence as an obstacle to deciding when retrieval is needed. Mitigating overconfidence improves self-assessment and helps the model seek external knowledge when its own knowledge is insufficient.
The result connects evaluation to action: better knowledge-boundary perception supports comparable or stronger question-answering performance with fewer retrieval calls. Self-assessment matters because it can improve decisions, not just evaluation scores.
Enlarge figure ↗Reducing confidence in incorrect answers helps expose the knowledge gaps that motivate retrieval. Figure from the ACL 2024 Findings paper.
Supporting evidence · The capability–confidence gap
We therefore look beyond a model’s verbal confidence to its internal representations. Our ACL 2025 study finds useful knowledge-boundary signals before answer generation and examines how generation changes them, opening the possibility of assessing knowledge gaps before producing a complete answer.
For post-generation assessment, Confidence Consistency-based Calibration uses reformulated questions to improve the recognition of those gaps. Together with our comparison of confidence expressions below, this work distinguishes the information a model contains from how reliably it communicates that information.
Enlarge figure ↗Probing internal representations before and after generation reveals when self-assessment signals become available. Figure from the ACL 2025 paper.
↓These studies establish the value of self-assessment and identify useful signals. The next challenge is to understand—and reduce—their errors.
Chapter 02
Explaining and improving self-assessment
The gap between capability and self-assessment raises two connected questions: what systematically biases confidence, and how can reliable confidence be learned beyond a single task?
To understand why self-assessment fails, we examine knowledge popularity. Popular wrong answers—and answers strongly associated with the question—can attract high confidence. This reveals a structured bias: familiarity can be mistaken for evidence of correctness.
Incorporating popularity-related signals improves calibration. The work thus connects an explanation of overconfidence to a practical way of mitigating it: understanding what biases confidence helps us correct its relationship to actual performance.
Enlarge figure ↗A familiar wrong alternative can attract high confidence. The paper distinguishes ground-truth and generation-side popularity signals.
↓Explaining a specific bias is one part of the problem. A complementary challenge is learning reliable confidence across tasks without labeling every answer.
EliCal addresses this learning problem by separating two objectives: eliciting an informative confidence signal and calibrating it against correctness. Large-scale self-consistency supervision supports the first stage; a small set of correctness labels supports the second. HonestyBench provides the training and evaluation resource.
Where popularity-aware calibration targets an identifiable bias, EliCal studies how to learn reliable confidence more generally. Its two-stage design reduces dependence on costly correctness labels, making annotation efficiency part of the honesty-alignment problem.
1,000correctness annotations
About 0.18% of the full-supervision correctness labels, approaching fully supervised alignment in the reported experiments; additional self-consistency supervision is used.
Enlarge figure ↗EliCal separates low-cost confidence elicitation from correctness-based calibration, with HonestyBench as the supporting data resource.
↓EliCal reduces the cost of learning honest confidence. But as task training changes what a model can do, its self-assessment must also keep pace.
Chapter 03
Learning confidence alongside capability
A model’s capabilities change during training. Rather than calibrating only the final model, can we also learn from the successes, failures, and verified outcomes along the way?
Our latest work, CoCal, extends the learning question to capability-building experience itself. A lightweight confidence companion learns from internal states and verifier feedback in rollouts collected during reinforcement learning with verifiable rewards (RLVR).
Task learning and confidence learning share experience, but use separate parameters and optimization; the task-training objective remains unchanged. This explores a further principle for the research program: the experience that builds capability can also support learning to judge that capability, without coupling the two objectives.
0.120 → 0.096expected calibration error ↓
Matched post-hoc calibration → CoCal (binary). Qwen3-8B, five math benchmarks, 120 training steps; both retain 63.8% task accuracy.
Enlarge figure ↗CoCal reuses capability-training experience without coupling confidence learning to policy optimization. Figure 1 from the CoCal preprint.
Scope. The companion requires access to hidden states and does not directly change verbalized confidence. Experiments train it by replaying cached training experience in chronological order.
What this program adds up to
Our question has developed from whether models recognize their limits, to why they misjudge them, and how that judgment can be learned as capabilities change. The goal is not indiscriminate caution, but self-assessment that supports better decisions. Confidence remains a signal—not proof of correctness.
03 / Testing the scope of this perspective
Extending to more demanding settings
Our related studies test this perspective as the evidence, the meaning of correctness, and the object of evaluation become more complex—from visual answers to factuality judgments and ranked evidence.
Carries self-assessment from answer correctness to the quality of ranked evidence.
A broader view of research agents. In Deep Research: A Systematic Survey, I contributed the sections on retrieval timing and its open challenges (§3.2.2 and §6.1), connecting knowledge-boundary perception to reliable information-seeking decisions.
04 / Research outlook
From “I may be wrong” to “here is what would help.”
From assessing an answer to diagnosing and improving it.
The program so far asks how reliably a model can assess its answers. In open-ended tasks, an answer may be only partly correct. Our next question is finer-grained: which claim is wrong or uncertain, why, and what evidence or reasoning would resolve the gap?
Such diagnoses could guide retrieval, clarification, and revision. Independently verified outcomes could then support further learning—connecting knowledge-boundary perception to iterative self-improvement.
LocateWhich claim is uncertain or wrong?
DiagnoseWhat knowledge, evidence, or reasoning is missing?
CorrectWhat action would resolve the gap?
Verify & learnDid the correction work, and does it transfer?
More capable models should also become better judges of their own limits.