Confidence is not all you need

A critique of self-confidence as a learning signal.

The promise of Confidence Is All You Need is appealing: let a model generate answers, reward the ones it considers likely, and improve reasoning without labels or an external judge.

I wanted that idea to work. My first encouraging result seemed to support it. Then I added controls, repeated the experiment, and inspected the actual learning signal.

My objection is simple: a better checkpoint does not establish that confidence caused the improvement. The paper reports gains, but I do not think it adequately separates that claim from the much stronger claim that self-confidence is a reliable learning signal.

I tested the published idea in an independent implementation using Qwen models. These are my results, not an audit of the authors' training runs or an exact replication of their entire evaluation setup.

Being predictable is not the same as being right

The paper's objective rewards concentration: the model puts more probability on responses it already favors. That objective has no direct way to distinguish a correct answer from a confidently repeated mistake.

Imagine a model that usually gives the wrong answer but occasionally finds the right one. Reinforcing its preferred response can make it more consistent while making the correct answer harder to discover. Whether confidence helps therefore depends on which responses the model already prefers.

There is also a length problem. Whole-response probability multiplies the probabilities of every token. A long correct solution can receive less weight than a short incorrect one. Calling that number “confidence” does not turn it into a probability of correctness.

This does not make confidence useless. It makes the missing question more important: when does rewarding likelihood actually improve correctness?

The control changes the story

I compared the original model with two trained copies: one configured to use confidence weighting, and one giving every sampled response the same weight. I call the first arm nominal RLSC, because the configured confidence term did not always affect the applied weights.

Both trained copies learned from their own generated answers. The equal-weight control tests whether adding confidence provides anything beyond that shared procedure.

Math-model accuracy on MATH-500 (%)
ExperimentOriginalNominal RLSCEqual weights
1.5B · 10 updates52.6046.2054.80
1.5B · 20 updates51.8043.8053.40
7B · 10 updates, run A58.6069.8062.40
7B · 20 updates58.0066.6075.00
7B · 10 updates, run B58.2060.8069.60

Green and bold mark the highest score in each row. Each run starts from the original model; run A and run B use different random seeds. Original scores are the recorded baselines for each setting.

The confidence-configured arm wins one of these five comparisons. Equal weighting wins the other four. The initial 7B improvement looked exciting, but its relative advantage did not survive a longer run or another seed. An audit also found constant applied weights in that early positive run.

General Qwen models were less reassuring: nominal RLSC finished below the original model in all six ten-update settings I tested. A lower learning rate rescued some performance, but still did not produce a consistent confidence advantage.

These results do not erase the paper's reported gains. They challenge the explanation. Its full smoothing ablation is deferred, and the reported benchmark comparison does not establish a benefit over equal-weight self-training. That control is central to the claim.

Sometimes confidence is not even in the update

The published recipe adds a constant to whole-response probability. I tested its suggested value, giving a weight of p(response) + 0.1.

Whole-response probabilities get tiny. Even an answer with 200 tokens, each assigned probability 0.9, has a sequence probability of about 0.0000000007. In the arithmetic used for my training weights, adding that to 0.1 changes nothing.

In one closely audited 14B training pair, all 320 nominal confidence weights were identical to the equal-weight control. Most probabilities became zero during exponentiation; the rest were too small to change the weight after adding 0.1.

That is a concrete failure mode of my implementation of the published rule. It is not evidence that the authors' actual runs had the same numerical behavior. But it shows why a confidence-based method needs to report the weights that actually reach training.

The two models can still end up different because they sampled different answers. Their different scores cannot demonstrate a confidence effect when confidence never changed their weights.

What happens to the answers the model rarely finds?

A model becoming more certain would not, by itself, convince me it had improved. I also want to know whether it can still find correct answers through sampling.

I checked this on 64 held-out questions across a 14B training pair. Pass@1 estimates the chance that one sampled answer is correct. Pass@16 asks whether sixteen attempts contain at least one correct answer.

The original model solved 90.6% of these questions within sixteen attempts. After training, that coverage fell to 43.8% for nominal RLSC and 79.7% for the equal-weight control. Single-sample accuracy also fell in both arms.

On eleven questions where correct answers were initially rare, fresh sampling recovered nine with the original model, five with the control, and none with nominal RLSC.

This was the pair with identical applied weights, so I cannot blame that difference on active confidence weighting. Many responses also ran out of tokens before finishing. The finding is narrower: useful answer coverage deteriorated under this setup. It does not prove the model permanently lost those reasoning paths.

Still, this is exactly the kind of test a mode-sharpening method needs. A headline accuracy gain cannot tell the whole story about useful exploration.

What would convince me?

Confidence helped rank existing answers in my tests. That is a useful result. Choosing among answers already generated, however, is a different claim from using confidence to improve the model that generates them.

For the training claim, I would want three things:

  • Evidence that confidence meaningfully changes the applied weights.
  • A repeatable advantage over both the original model and equal-weight self-training.
  • Checks that correct answers remain recoverable through sampling, with response length and unfinished answers accounted for.

My experiments do not show that every confidence-based method must fail. They show why the title asks me to accept too much. Confidence can be informative, numerically irrelevant, or reinforce an existing mistake. A convincing method has to establish which one is happening.

I came looking for self-improvement. I found that the first thing to verify is whether the supposed learning signal is doing anything at all.

The supporting results and protocol contain the full comparisons. Scores here use my evaluation setup; they are not presented as direct numerical reproductions of the paper's reported scores.