Shiyu NiResearch

A connected research program · Shiyu Ni · 2024–2026

Knowledge boundaries.
Honest AI.

Knowing an answer is not the same as knowing when to trust it.

Our work connects knowledge-boundary evaluation, the explanation of overconfidence, and honesty alignment—with one goal: helping models make better decisions about their own answers.

Explore the research story

01 / Our perspective

Self-assessment is part of decision-making.

Recognizing its limits helps a model decide when to answer, seek evidence, or defer. We study how to make this judgment reliable—not merely how to attach a confidence score.

Five-part research framework. Definition: knowledge-boundary perception means knowing what the model knows and what it does not know. The blue region represents what the model knows; the red region represents what it thinks it knows. Blue-only areas indicate underconfidence; red-only areas indicate overconfidence. Evaluation: calibration and error discrimination. Explanation: popularity is associated with confidence but is not correctness. Improvement: calibrate signals, perform honesty alignment with fewer labels, and learn confidence alongside capability. Applications: answer or abstain, adaptive retrieval, model routing, and reasoning budgets, taking task needs and costs into account. Verified feedback for self-improvement is marked as an outlook. All plots are conceptual illustrations.
A conceptual overview of the research direction; plots are schematic. Click to read at full size.Open original

02 / Developing the argument

From observing a gap
to learning reliable self-assessment.

Across this program, one question persists: how can a model’s judgment of its own answers better reflect what it can actually do?

Chapter 01

Understanding knowledge boundaries

We began with a practical question: when should a model rely on its own knowledge, and when should it retrieve? Answering it requires measuring the gap between what a model knows and what it thinks it knows.

When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation

ACL 2024 Findings First author

Our ACL 2024 Findings study makes this gap measurable and identifies overconfidence as an obstacle to deciding when retrieval is needed. Mitigating overconfidence improves self-assessment and helps the model seek external knowledge when its own knowledge is insufficient.

The result connects evaluation to action: better knowledge-boundary perception supports comparable or stronger question-answering performance with fewer retrieval calls. Self-assessment matters because it can improve decisions, not just evaluation scores.

Paper figure comparing confidence in incorrect answers before and after confidence alignment across models and datasets.Enlarge figure
Reducing confidence in incorrect answers helps expose the knowledge gaps that motivate retrieval. Figure from the ACL 2024 Findings paper.

Supporting evidence · The capability–confidence gap

Knowing vs. thinking you know. Natural Questions · vanilla prompting · models evaluated in the 2024 study. Metric: Share of answers (%). Values and experimental notes are available below.Enlarge chart
Redrawn from ACL 2024 Findings · Table 2, p. 4; definitions in Section 3.2.SVGPNGData
Data and experimental setting

Natural Questions · vanilla prompting · models evaluated in the 2024 study. Share of answers (%).

‘Certain’ is a binary self-report, not an average probability. Gap = certain rate − accuracy (percentage points).

This aggregate gap is not the paper’s sample-level Overconfidence metric. All five Table 2 models are shown.

Knowing vs. thinking you know
SettingActually correctSays ‘certain’
Vicuna26.3497.22
LLaMA229.8683.16
GPT-Instruct40.0381.0
ChatGPT38.570.83
GPT-448.9681.06

Once confidence guides decisions, its source matters: does the model reveal everything it knows about its own limits?

Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception

ACL 2025 First author

We therefore look beyond a model’s verbal confidence to its internal representations. Our ACL 2025 study finds useful knowledge-boundary signals before answer generation and examines how generation changes them, opening the possibility of assessing knowledge gaps before producing a complete answer.

For post-generation assessment, Confidence Consistency-based Calibration uses reformulated questions to improve the recognition of those gaps. Together with our comparison of confidence expressions below, this work distinguishes the information a model contains from how reliably it communicates that information.

Paper diagram showing where internal representations are extracted before and after answer generation.Enlarge figure
Probing internal representations before and after generation reveals when self-assessment signals become available. Figure from the ACL 2025 paper.

These studies establish the value of self-assessment and identify useful signals. The next challenge is to understand—and reduce—their errors.

Chapter 02

Explaining and improving self-assessment

The gap between capability and self-assessment raises two connected questions: what systematically biases confidence, and how can reliable confidence be learned beyond a single task?

Popular but Wrong: Understanding and Mitigating LLM Overconfidence through Knowledge Popularity

EMNLP 2026 Oral First author

To understand why self-assessment fails, we examine knowledge popularity. Popular wrong answers—and answers strongly associated with the question—can attract high confidence. This reveals a structured bias: familiarity can be mistaken for evidence of correctness.

Incorporating popularity-related signals improves calibration. The work thus connects an explanation of overconfidence to a practical way of mitigating it: understanding what biases confidence helps us correct its relationship to actual performance.

Knowledge-popularity framework contrasting a popular wrong answer with a correct answer and illustrating popularity-aware confidence calibration.Enlarge figure
A familiar wrong alternative can attract high confidence. The paper distinguishes ground-truth and generation-side popularity signals.

Explaining a specific bias is one part of the problem. A complementary challenge is learning reliable confidence across tasks without labeling every answer.

Annotation-Efficient Universal Honesty Alignment

ICLR 2026 First author

EliCal addresses this learning problem by separating two objectives: eliciting an informative confidence signal and calibrating it against correctness. Large-scale self-consistency supervision supports the first stage; a small set of correctness labels supports the second. HonestyBench provides the training and evaluation resource.

Where popularity-aware calibration targets an identifiable bias, EliCal studies how to learn reliable confidence more generally. Its two-stage design reduces dependence on costly correctness labels, making annotation efficiency part of the honesty-alignment problem.

1,000correctness annotations

About 0.18% of the full-supervision correctness labels, approaching fully supervised alignment in the reported experiments; additional self-consistency supervision is used.

EliCal two-stage training diagram: confidence elicitation from large-scale consistency data, followed by calibration using a small correctness-labeled set.Enlarge figure
EliCal separates low-cost confidence elicitation from correctness-based calibration, with HonestyBench as the supporting data resource.

EliCal reduces the cost of learning honest confidence. But as task training changes what a model can do, its self-assessment must also keep pace.

Chapter 03

Learning confidence alongside capability

A model’s capabilities change during training. Rather than calibrating only the final model, can we also learn from the successes, failures, and verified outcomes along the way?

Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs

Preprint First authorUnder review

Our latest work, CoCal, extends the learning question to capability-building experience itself. A lightweight confidence companion learns from internal states and verifier feedback in rollouts collected during reinforcement learning with verifiable rewards (RLVR).

Task learning and confidence learning share experience, but use separate parameters and optimization; the task-training objective remains unchanged. This explores a further principle for the research program: the experience that builds capability can also support learning to judge that capability, without coupling the two objectives.

0.120 → 0.096expected calibration error ↓

Matched post-hoc calibration → CoCal (binary). Qwen3-8B, five math benchmarks, 120 training steps; both retain 63.8% task accuracy.

CoCal conceptual comparison: post-hoc calibration, coupled reinforcement learning, and shared experience with separate capability and confidence learning.Enlarge figure
CoCal reuses capability-training experience without coupling confidence learning to policy optimization. Figure 1 from the CoCal preprint.

Scope. The companion requires access to hidden states and does not directly change verbalized confidence. Experiments train it by replaying cached training experience in chronological order.

What this program adds up to

Our question has developed from whether models recognize their limits, to why they misjudge them, and how that judgment can be learned as capabilities change. The goal is not indiscriminate caution, but self-assessment that supports better decisions. Confidence remains a signal—not proof of correctness.

03 / Testing the scope of this perspective

Extending to more demanding settings

Our related studies test this perspective as the evidence, the meaning of correctness, and the object of evaluation become more complex—from visual answers to factuality judgments and ranked evidence.

A broader view of research agents. In Deep Research: A Systematic Survey, I contributed the sections on retrieval timing and its open challenges (§3.2.2 and §6.1), connecting knowledge-boundary perception to reliable information-seeking decisions.

04 / Research outlook

From “I may be wrong”
to “here is what would help.”

From assessing an answer to diagnosing and improving it.

The program so far asks how reliably a model can assess its answers. In open-ended tasks, an answer may be only partly correct. Our next question is finer-grained: which claim is wrong or uncertain, why, and what evidence or reasoning would resolve the gap?

Such diagnoses could guide retrieval, clarification, and revision. Independently verified outcomes could then support further learning—connecting knowledge-boundary perception to iterative self-improvement.

  1. LocateWhich claim is uncertain or wrong?
  2. DiagnoseWhat knowledge, evidence, or reasoning is missing?
  3. CorrectWhat action would resolve the gap?
  4. Verify & learnDid the correction work, and does it transfer?

More capable models should also become
better judges of their own limits.

Browse all publications