Thursday, October 1, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Evening Edition Free of Charge
TODAY'S LEAD STORY

Testing the "Hallucination is due to the free version" theory with benchmarks—how the answer can be reversed depending on the measurement method.

This study verifies the claim on X that "hallucination is no longer a topic of discussion because it's only an issue for free version users," using data from two benchmarks: Vectara HHEM and AA-Omniscience. In contrast to contextual summarization tasks, flagship models exhibit a higher hallucination rate in uncontextual knowledge recall tasks, demonstrating a "non-abstention" structure, which is explored through the case of Kimi K3.

Testing the "Hallucination is due to the free version" theory with benchmarks—how the answer can be reversed depending on the measurement method.
(Photo: illustrative)

A Certain Kind of "Reassurance" Seen on X

"The reason we don't see much talk about hallucination lately is because more people are paying for and using the flagship model; the ones still complaining are those using the free version." Posts to this effect have been trending on X recently. Intuitively, this sounds logical. Indeed, OpenAI itself explains that when comparing their free ChatGPT model, "GPT-5.6 Luna," released in July, with their paid model, "GPT-5.6 Sol," Sol reduces the hallucination rate by 68%. It's true that there is a real performance difference between the free and paid versions. However, concluding that "hallucination is now only a problem for the free version" based solely on this one figure is a rather dangerous simplification when you examine the benchmarks.

Benchmarks Can Give You the Opposite "Answer"

This is the most important point I want to make here. While hallucination rate appears to be a single metric, its evaluation actually varies greatly depending on the measurement method. In Vectara's summarization task benchmark (which measures the accuracy of summarizing a given text), Anthropic, OpenAI, and Google's frontier models are projected to have hallucination rates of approximately 7-15% by 2026. Flagship models generally perform well in summarization tasks where context (source documents) is provided. However, the situation changes dramatically in "AA-Omniscience," a benchmark developed by Artificial Analysis that asks users to answer difficult fact-checking questions without context. GPT-5.6 Sol achieved a hallucination rate of 92.2%, and Claude Fable 5 achieved 63.6%, both higher than some lesser-known, smaller models.

The Price of the "No Answer" Option

Why do the results vary so drastically? The reason lies in AA-Omniscience's scoring method. This benchmark examines which of three options a model chooses: "correct answer," "incorrect answer," or "don't know and withdraw." The hallucination rate is calculated by dividing the number of incorrect answers by the total number of incorrect answers plus withdrawals. In other words, models that honestly answer "I don't know" (withdraw) when they don't know the answer will have a lower hallucination rate. Conversely, models that confidently attempt to answer even unknown questions will often be wrong, resulting in a high hallucination rate. Artificial Analysis's own analysis frankly acknowledges this point—"There is a correlation between model size and accuracy, but not with hallucination rate." This means that smarter, more expensive models tend to attempt to answer even unknown questions, resulting in a counterintuitive structure where hallucination rates tend to be higher in this type of benchmark.

The "Increased Accuracy, Increased Hallucination Rate" Phenomenon as Demonstrated by KimiK3

A prime example that succinctly illustrates this structure is Moonshot AI's Kimi K3. The upgrade from the previous model's K2.6 to K3 significantly improved the accuracy rate on AA-Omniscience from 33% to 46%. However, at the same time, the hallucination rate worsened from 39% to 51%. While the overall score (Omniscience Index) itself improved from +6 to +18 during this period, this is because the scoring system is designed to give more weight to improvements in accuracy rate than to deterioration in hallucination rate. These figures clearly demonstrate that "becoming smarter" and "reducing hallucination" are not necessarily changes that go in the same direction.

The Partial Correctness of "Paying for Peace of Mind"

Based on the above, let's distinguish between what is correct and what is questionable in X's initial claim. In many everyday use cases, such as summarizing, organizing information, and coding assistance given context, the flagship model tends to produce more consistently stable output than the free, lightweight model, and in this sense, the intuition that "paying for peace of mind" is not entirely off the mark. On the other hand, in terms of a more fundamental axis of reliability—whether the model can honestly acknowledge the limitations of its own knowledge and admit when it doesn't know something—flagship models are not necessarily superior, and in some cases, they even perform worse. The simplistic pitfall inherent in the discussion at X is the failure to distinguish between these two axes and simply assuming "flagship models are reliable."

What Researchers Should Consider

Behind the single number of "hallucination rate" lies a complex interplay of the measurement task design, scoring method, and the model's strategic choice of "answering/not answering." Judging a model as reliable or unreliable based solely on a single benchmark score is premature; at the very least, it's necessary to consciously differentiate between two distinct evaluation axes: "contextual summarization tasks" and "uncontextual knowledge recall tasks." We will continue to closely monitor how model development companies' design decisions regarding whether to target "intelligence" or "honest abstention" for optimization will narrow the gap between these kinds of benchmarks.

ハルシネーションAA-OmniscienceベンチマークLLM評価GPT-5.6

An agent that "continues to work even after the conversation ends"—the essence of "Dots" revealed at OpenAI DevDay 2026

OpenAI announced "Dots," a new agent that operates independently of conversations and runs continuously, at DevDay 2026 on September 29th. This article will analyze its features, including its GPT-6.1 Sol architecture (which runs at one-fifth the cost of GPT-6 Astra), its integration with 4000 applications, its secure design with approval gates and private safety processing, and its competitive landscape with Meta's Muse.

An Agent That Continues Working Even After the Conversation Ends

On September 29th, at OpenAI DevDay 2026 in San Francisco, CEO Sam Altman announced a new agent called "Dots." Previous ChatGPT systems were based on a question-and-answer structure, where the user asked a question, and the conversation ended once an answer was received. Dots overturns this premise. It operates independently of conversations, constantly running, possessing its own dedicated cloud computer and browser, and continuously handling workflows ranging from software development and data analysis to internal operations through integration with 4000 types of applications.

Design from Both a Foundational and Cost Perspective

Dots is powered by OpenAI's top-of-the-line model, GPT-6 Astra. The simultaneously announced GPT-6.1 Sol is a highly efficient model optimized for agent-based coding, offered at one-fifth the computational cost of Astra. At DevDay, a "hosted browser desktop" was also offered via the new Agents API, allowing developers to use a virtual desktop environment where agents can click and type without having to build their own. While there were flashy feature additions, the fact that cost efficiency—a subtle but crucial aspect—was also addressed is commendable.

Integration of "Specialized Dots" with Microsoft

In addition to Dots for individuals, "Specialist Dots" for enterprises was also announced. These are agents with permissions and functions narrowed down for specific business roles, designed to integrate with Microsoft Agent 365 and be governed through existing management tools. The biggest concerns for companies when deploying AI agents are permission management and auditability. The decision to integrate with the existing enterprise management platform is a realistic approach that looks ahead to actual operation.

How to balance "greater autonomy" and "security"

For always-on agents like Dots, the more autonomy they have, the greater the risk. OpenAI addresses this by combining an approval gate (a mechanism that requires user approval before an agent takes certain actions) with a feature called "private safety processing." The latter allows for automated safety reviews without creating a new channel for OpenAI personnel to directly access protected customer content. Dots is said to perform "proactive research" even when users are not actively working on tasks, through a read-only tool they connect to. With each such feature enhancement, it's necessary to examine how the scope of permissions and privacy protection are balanced.

Context of Competition with Meta's Muse

The timing of the Dots announcement comes shortly after Meta's personal AI agent, "Muse," garnered significant attention. While Muse positions itself as an "agent-type companion" for general consumers, OpenAI positions Dots as a more business-oriented "always-on executioner." The fact that various companies are exploring the same "personal agent" concept, but on different axes—consumer-oriented and business-oriented—is emerging as a major trend this fall.

Things Engineers Should Consider

The concept of an "agent that continues working even between conversations" is an extension of the design philosophy of "continuously advancing tasks in the background" that has been cultivated in previous AI coding agents. However, what's new about Dots is that it aims to apply this design to a wider range of business areas, such as schedule management and internal research, beyond the limited domain of coding. It is planned to start with availability to Pro and Business Premium users and then gradually expand to opt-in beta for Enterprise, Education, and Healthcare users. The success or failure of Dots will likely depend on how many users are willing to entrust "meaningful work" to this type of agent and let it go.

OpenAIDotsAIエージェントDevDayGPT-6

Which is correct: a $50 trillion market or a $32.4 trillion market? — An accountant analyzes SiMa.ai's $1.45 billion valuation.

On September 28, SiMa.ai, a semiconductor company specializing in physical AI, announced it had raised $150 million in Series C funding, bringing its valuation to $1.45 billion. This article examines, from an accountant's perspective, the inconsistencies in the company's own market size figures, the fact that its next-generation chip, scheduled for release in 2028, will only offer half the performance of current NVIDIA products, and its differentiation strategy of focusing on "power and cost rather than performance."

Which is Correct: A $50 Trillion Market or a $32.4 Trillion Market?

On September 28th, SiMa.ai, a company developing semiconductors for physical AI, announced that it had raised $150 million in Series C funding, bringing its valuation to $1.45 billion. The round was co-led by Fidelity Management & Research Company and Amplify, with participation from multiple investors including Dell Technologies Capital and Point72. The total amount raised has reached $500 million, representing an approximately 51% increase in just over a year from its Series B valuation of $960 million in July 2025. As an accountant, what caught my eye wasn't so much the valuation trend itself, but rather a slight inconsistency in the company's use of "market size" figures.

Market Size as Described by the CEO: Figures Conflict in Different Reports

Founder and CEO Krishna Rangasai describes physical AI in humanoid robots, automobiles, and drones as "the gateway to a $50 trillion market that has remained largely untouched by modern innovation." However, other reports use the figure of "$32.4 trillion market" in the same context. It's important to note that neither figure is explicitly backed by independent market research from a third-party organization; these are merely the company's own stated estimates of the future market size. While these kinds of "trillion-dollar" market size forecasts can raise investor expectations for startups, they are also prone to being misleading and easily misinterpreted. The very existence of two figures with vastly different magnitudes clearly illustrates the nature of this type of market size forecast.

Technical Positioning – Competing on "Power and Cost" Rather Than "Performance"

SiMa.ai offers a full-stack platform combining dedicated silicon and software for directly running AI models on physical devices such as robots, autonomous vehicles, and drones. The funds raised will be used to expand the development environment "Palette Neat" and to develop the next-generation system-on-a-chip, "Modalix." From an accounting perspective, the performance target for this next-generation chip is interesting. SiMa.ai aims to launch its new chip in the first half of 2028 with a performance of 1000 TOPS (tera-computation per second). Meanwhile, NVIDIA's Jetson AGX Thor is said to already achieve 2000 FP4 TOPS. This means that even SiMa.ai's next-generation chip, which it aims to launch in two years, will only have half the computing performance of NVIDIA's current product. The company recognizes this and has clearly emphasized differentiation not on "raw computing performance," but on a different axis: "low power consumption and low cost."

The Significance of a Top 5% Funding Scale

The $150 million funding round itself needs to be evaluated within its context. One analysis points out that this amount places the company in the top 5% of Series C funding rounds in the US enterprise software sector. The fact that a physical AI company, which combines hardware and software, has reached a funding scale that puts it on par with pure software companies is one indicator of the high expectations investors have for this field. Counterpoint Research predicts that the cumulative number of physical AI devices, including robotics, automobiles, and drones, will reach 145 million units by 2035, and this prediction is one of the supporting factors for this funding round.

A Bet to Break Free from NVIDIA Dependence

Another significance of this funding round is that investors continue to invest in the bet that "not all AI workloads run on NVIDIA GPUs." Unlike learning and inference for data centers, AI execution on edge devices such as robots and drones is heavily constrained by power consumption, thermal design, and cost. In this area, investment in specialized chip companies like SiMa.ai continues, based on the view that NVIDIA's general-purpose GPU architecture is not necessarily the optimal solution.

Points to Note from an Accountant's Perspective

The discrepancy in market size figures between "$50 trillion" and "$32.4 trillion" may seem trivial, but it symbolizes how the funding stories of this type of startup are often built on future predictions that are difficult to verify. The fact that the next-generation chip, scheduled for release in 2028, will not even reach the performance of NVIDIA's current products at launch is a reflection of this company's more modest but solid strategy of "winning on niche power and cost requirements" rather than "winning on performance." Going forward, we should closely monitor whether Modalix's development progresses as planned and how many orders it actually secures before the next funding round.

SiMa.aiフィジカルAI半導体資金調達NVIDIA
Advertisement300 × 250