A Certain Kind of "Reassurance" Seen on X
"The reason we don't see much talk about hallucination lately is because more people are paying for and using the flagship model; the ones still complaining are those using the free version." Posts to this effect have been trending on X recently. Intuitively, this sounds logical. Indeed, OpenAI itself explains that when comparing their free ChatGPT model, "GPT-5.6 Luna," released in July, with their paid model, "GPT-5.6 Sol," Sol reduces the hallucination rate by 68%. It's true that there is a real performance difference between the free and paid versions. However, concluding that "hallucination is now only a problem for the free version" based solely on this one figure is a rather dangerous simplification when you examine the benchmarks.
Benchmarks Can Give You the Opposite "Answer"
This is the most important point I want to make here. While hallucination rate appears to be a single metric, its evaluation actually varies greatly depending on the measurement method. In Vectara's summarization task benchmark (which measures the accuracy of summarizing a given text), Anthropic, OpenAI, and Google's frontier models are projected to have hallucination rates of approximately 7-15% by 2026. Flagship models generally perform well in summarization tasks where context (source documents) is provided. However, the situation changes dramatically in "AA-Omniscience," a benchmark developed by Artificial Analysis that asks users to answer difficult fact-checking questions without context. GPT-5.6 Sol achieved a hallucination rate of 92.2%, and Claude Fable 5 achieved 63.6%, both higher than some lesser-known, smaller models.
The Price of the "No Answer" Option
Why do the results vary so drastically? The reason lies in AA-Omniscience's scoring method. This benchmark examines which of three options a model chooses: "correct answer," "incorrect answer," or "don't know and withdraw." The hallucination rate is calculated by dividing the number of incorrect answers by the total number of incorrect answers plus withdrawals. In other words, models that honestly answer "I don't know" (withdraw) when they don't know the answer will have a lower hallucination rate. Conversely, models that confidently attempt to answer even unknown questions will often be wrong, resulting in a high hallucination rate. Artificial Analysis's own analysis frankly acknowledges this point—"There is a correlation between model size and accuracy, but not with hallucination rate." This means that smarter, more expensive models tend to attempt to answer even unknown questions, resulting in a counterintuitive structure where hallucination rates tend to be higher in this type of benchmark.
The "Increased Accuracy, Increased Hallucination Rate" Phenomenon as Demonstrated by KimiK3
A prime example that succinctly illustrates this structure is Moonshot AI's Kimi K3. The upgrade from the previous model's K2.6 to K3 significantly improved the accuracy rate on AA-Omniscience from 33% to 46%. However, at the same time, the hallucination rate worsened from 39% to 51%. While the overall score (Omniscience Index) itself improved from +6 to +18 during this period, this is because the scoring system is designed to give more weight to improvements in accuracy rate than to deterioration in hallucination rate. These figures clearly demonstrate that "becoming smarter" and "reducing hallucination" are not necessarily changes that go in the same direction.
The Partial Correctness of "Paying for Peace of Mind"
Based on the above, let's distinguish between what is correct and what is questionable in X's initial claim. In many everyday use cases, such as summarizing, organizing information, and coding assistance given context, the flagship model tends to produce more consistently stable output than the free, lightweight model, and in this sense, the intuition that "paying for peace of mind" is not entirely off the mark. On the other hand, in terms of a more fundamental axis of reliability—whether the model can honestly acknowledge the limitations of its own knowledge and admit when it doesn't know something—flagship models are not necessarily superior, and in some cases, they even perform worse. The simplistic pitfall inherent in the discussion at X is the failure to distinguish between these two axes and simply assuming "flagship models are reliable."
What Researchers Should Consider
Behind the single number of "hallucination rate" lies a complex interplay of the measurement task design, scoring method, and the model's strategic choice of "answering/not answering." Judging a model as reliable or unreliable based solely on a single benchmark score is premature; at the very least, it's necessary to consciously differentiate between two distinct evaluation axes: "contextual summarization tasks" and "uncontextual knowledge recall tasks." We will continue to closely monitor how model development companies' design decisions regarding whether to target "intelligence" or "honest abstention" for optimization will narrow the gap between these kinds of benchmarks.