Monday, July 13, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Evening Edition Free of Charge
TODAY'S LEAD STORY

AI "bound itself to avoid hallucinations"—the ingenuity of an autonomous research agent tackling an unsolved problem in physics

This article explains the paper "Grounded autonomous research" published by Haonan Huang on July 2nd. It describes a pipeline that autonomously selects research topics from 11,083 arXiv papers in condensed matter physics, calibrates methods through numerical comparison with previously published papers, and then performs new calculations and writes papers. The article deciphers the design philosophy of "grounding in topic selection," which involves preemptively discarding topics for which there is no relevant prior research, in order to prevent hallucinations.

AI "bound itself to avoid hallucinations"—the ingenuity of an autonomous research agent tackling an unsolved problem in physics
(Photo: illustrative)

AI Binds Itself to Avoid Hallucinations—An Autonomous Research Agent's Strategy in Tackling an Unsolved Problem in Physics

Previously in this column, I introduced Sakana AI et al.'s "The AI ​​Scientist," which was published in Nature and whose generative paper passed peer review at the ICLR workshop. That was in the field of machine learning, but the paper I'll introduce this time tackles a much more challenging area—cutting-edge physics—with an autonomous research agent. Let's look at the single-authored paper "Grounded autonomous research" by Haonan Huang, published on arXiv on July 2nd.

Why is Physics More Difficult Than a "Machine Learning Sandbox"?

This paper first carefully points out the essential difference between the areas demonstrated by AI autonomous research agents to date and physical science. In machine learning experiments (writing and running code and measuring accuracy), the execution result itself becomes the criterion for determining correctness (calibration). When code runs, numbers are produced, and those numbers themselves can serve as evidence of "correctness."

However, cutting-edge physics is different. Every methodological choice requires physical reasoning, the toolchains used are often not adequately documented, and above all, the criteria for verifying the validity of results exist only in external literature. This paper points out a dangerous weakness inherent in existing autonomous agents. Existing agents cite external literature but do not actually confront its content with their own results, leading to the problem of "hallucination" of plausible but unverifiable results based solely on the AI ​​model's internal knowledge.

Choosing a Theme from 11,083 Papers

To summarize what this pipeline did, it focused on "altermagnetism," a relatively new field of condensed matter physics related to the classification of magnetism.

1. The agent autonomously maps a corpus of 11,083 recent arXiv papers in condensed matter physics to formulate a research direction. 2. By actually reproducing the numerical results of previously published papers, the agent calibrates the computational methodology it uses. 3. Using the calibrated methodology, it performs novel first-principles calculations that no one has ever performed before. 4. Based on the results obtained, it writes a draft of the paper.

What's interesting here is the theme selection process. The agent ultimately narrowed down five candidate themes to one: "altermagnetic piezomagnetism (an alter-magnet version of the phenomenon where magnetization occurs due to pressure)." According to an independent post-hoc review, this choice was sound. The other four candidate themes either required the development of new methods, relied on young toolchains not yet sufficiently validated for the target system, or had their novelty limited by existing research. In contrast, the chosen piezomagnetism theme had three independent, previously published papers available as numerical references for its target observable quantity.

This decision to select a theme based on the availability of relevant prior research is the core ingenuity of this paper. If they had ventured into an unverifiable theme, the pipeline itself would be unable to score the correctness of its own output—in other words, "grounding at the topic selection stage" is a prerequisite for grounding at the execution stage—a point that resonated deeply with me as an AI researcher.

A painstaking design: a relay of 47 "blank memories"

I would also like to touch upon the scale of the pipeline. The whole is divided into six phases and consists of 47 fresh context sessions (i.e., sessions starting from a blank slate, without any past conversation history). Each session proceeded in a relay format, sharing only the state saved on disk, and a total of 2,162 events involved referencing literature.

The design, which deliberately connects dozens of sessions without memory rather than dragging a long context into a single session, is a logical countermeasure to the phenomenon of "inference degrading as the context lengthens," a common problem in LLM agent practice. Rather, it is in this unassuming and down-to-earth design decision that the authors' practical knowledge shines through.

"Building a Reliable Footprint" Rather Than "Discovery"

Ultimately, this paper argues not for a flashy story of "AI making a new discovery," but rather for a design theory on how to create a safe footprint when entrusting research to AI agents in areas with high verification costs, such as physical science.

The previously introduced "The AI ​​Scientist" demonstrated its capabilities by actually passing workshop peer review in the relatively verifiable field of machine learning. Huang's paper goes a step further—demonstrating the possibility of autonomous research, even in more difficult-to-verify areas where physical reasoning is essential, by incorporating "confrontation" with external literature into the mechanism, while suppressing hallucinations. However, it's important to note that this is a single-authored paper, not yet formally published after peer review, and is a case study focusing on a single physical phenomenon. Therefore, the results should be taken with caution when generalizing them.

What concerns me as an AI researcher

To be honest, what I find most interesting in this paper is not the technical novelty itself, but the idea of ​​humans pre-designing what AI won't verify. The decision to discard four out of five candidate themes due to a lack of available prior research may seem conservative at first glance, but I believe it's an essential safety measure for ensuring the reliability of autonomous agents.

The theme of automating scientific research with AI is being explored from various angles as we enter 2026. The accumulation of meticulous design efforts like this one, aimed at ensuring verifiability, may determine whether this field transcends mere hype and becomes established as a truly reliable scientific method. I would like to follow up on further reports to see how far this method can be generalized to other fields and topics in physics.

AI/ML論文AI研究自動化物理学凝縮系物理arXiv

A robotic hand that can grasp a wine glass without breaking it—the 1X's innovative approach of a "low gear ratio."

This article provides a technical explanation of the new hand (25 degrees of freedom, tendon-driven, IP68 waterproof) for the NEO humanoid robot, announced by 1X Technologies on July 9th. It introduces the design philosophy of achieving backdrivability through a low gear ratio, about one-tenth that of industrial models, the idea that "the hand is a stack of perception," and the plan to mass-produce 10,000 units by 2026, while also touching on the limitations of self-reported specifications with a skeptical perspective.

A Robot Hand That Can Grasp a Wine Glass Without Breaking It—1X's Revolutionary Concept of "Low Gear Ratio"

On July 9th, 1X Technologies, the developer of the home humanoid robot "NEO," unveiled the full details of its new robot hand. 25 degrees of freedom (DoF), tendon drive, IP68 waterproofing—at first glance, the spec sheet seems like a string of numbers, but upon closer examination, it's packed with a paradigm shift in robot hand design, which was quite exciting for a software engineer like myself. This time, I'd like to break it down.

The Era of Competing on "Number of Degrees of Freedom" is Coming to an End

Until now, the competition surrounding humanoid robot hands has been solely based on the number of degrees of freedom. The Shadow Hand, which reigned as a research benchmark for over 20 years, had 24 joints. AGILINK, a subsidiary of China's AGIBOT, boasts over 8,000 units shipped with 20 degrees of freedom. At ICRA 2026 in June of this year, the "Wuji Hand 2," which pursues biomimetic and direct drive, was also unveiled.

1X's NEO hand has 25 degrees of freedom, which, in a simple numerical comparison, is only "slightly more than other companies." However, what 1X is emphasizing this time is not the number of degrees of freedom itself, but the fact that "all 25 joints are not passive spring joints, but are all actively force-controlled." In an industry where inflated numbers are rampant, the claim that they focused on "degrees of freedom that actually do work" rather than "degrees of freedom that are easy to count" is compelling.

The core is the reverse idea of ​​"reducing the gear ratio"

The most interesting aspect from a technical standpoint is the choice of drive system. 1X's NEO hand employs a pseudo-direct drive system where the motor is placed in the forearm (inside the arm, not the palm) and the fingers are moved through tendons. The gear ratio here is said to be around 5:1 to 15:1, which is an order of magnitude smaller than the typical 100:1 to 200:1 ratio used in industrial actuators.

What happens when the gear ratio is lowered? While high-gear-ratio actuators can greatly amplify the motor's rotational force, they also make it more difficult for external forces to be transmitted to the motor (low backdrivability). Conversely, lowering the gear ratio allows the change in force when the fingertips touch something to be transmitted directly to the motor through the tendons. To borrow the explanation from 1X, this design allows the hand to respond flexibly without resistance when encountering an unexpected object.

This is a subtle but essential point in the world of robot control. Conventional high-gear-ratio grippers have strong gripping power, but tend to struggle with the "control" needed to handle delicate objects without damaging them. The shift to a lower gear ratio can be understood as a design decision to restore the tactile feedback loop, even at the expense of power.

The phrase "The hand is a perception stack" is intriguing

Of all the 1X presentation materials, the phrase that personally resonated with me the most was, "The hand is not an actuator, but a perception stack." When humans touch an object, they press to check its hardness, slide their fingers to check its texture, and lift it to check its weight—in other words, touch is not a passive sensor, but rather the hand itself is an experimental device that actively "asks questions and reads answers".

This is a valid point from a cognitive science perspective, and it aligns with the process by which babies acquire manipulative abilities by repeatedly grasping and examining objects. 1X aims to design a system where this mechanism of "asking questions with force and reading answers with backdrivability" functions through the transmission mechanism itself, without the need for additional sensors.

The demo content is understated, but its understated nature is trustworthy

I also want to see the content of the publicly released demo. Assembling Lego blocks, picking up coins from a wallet, screwing in a light bulb, picking up screws, holding a wine glass, plugging in a USB-C cable, sorting grapes by color, lifting a 20-pound (approximately 9kg) kettlebell—the fact that it's a collection of everyday household tasks, rather than flashy somersaults or dances, is rather appealing.

However, I, Takahashi, want to emphasize this point. All the figures released so far are from 1X itself, and no verification or teardown (disassembly and inspection) by an independent third-party organization has yet been conducted. It's also unclear from current reports how much of the demo footage is autonomous operation and how much is remotely controlled (teleoperation). The 25 degrees of freedom figure itself isn't a dramatic difference compared to competitors like AGILINK (20 degrees of freedom) and Shadow Hand (24 degrees of freedom), and 1X's claim of being the "most advanced humanoid robotic hand" is merely self-proclaimed, which should be viewed calmly.

Another Significant Aspect: Mass Production

More than the technical aspects, what deserves attention is 1X's goal of manufacturing 10,000 hands by 2026, with in-house integrated production of everything from motors and tendon materials to polymers. While competitor AGILINK has already shipped 8,000 units, 1X is clearly ready to directly participate in this mass production race.

The observation that the competitive axis in the humanoid industry is shifting from "can it walk?" to "what can it do with its hands?" is spot on. A robot with only a simple two-fingered gripper can only offer developers three verbs: "grasp, place, and push." ​​On the other hand, a hand equipped with force control and haptic feedback offers a vastly expanded range of application possibilities.

Summary: "Reproducibility of Mass Production" Beyond the Spec Competition

The NEO hand's design philosophy is sound. However, what truly matters in this industry is not the "specs announced at the time of release," but the reproducibility of mass production—whether 10,000 units can actually be manufactured and each one can continue to function without breaking down. Pre-orders at the early access price of $20,000 have already begun, but whether the hands that actually reach consumers can maintain the same smoothness as the demo shown at the announcement—we will have to wait for future hands-on reviews to make a judgment.

1X TechnologiesNEOロボットハンドヒューマノイドフィジカルAIロボットコントローラ

"I'm not in acquisition negotiations"—Jim Keller himself denies the $1 billion deal rumors, but what's the truth?

Qualcomm CEO Jim Keller has denied speculation that the company is acquiring AI chip startup Tensorrent ($8 billion to $10 billion). This article examines the importance of distinguishing between confirmed information and speculative articles, comparing it to the recently confirmed acquisition of Modular (approximately $3.9 billion), the technical background of the RISC-V-based inference-focused architecture, and the concerns raised about talent acquisition risks.

"No Acquisition Negotiations Underway"—Jim Keller Himself Denies Billion-Dollar Rumors: What's the Truth?

Silicon Valley rumors are sometimes quickly dismissed by the person involved. This is a prime example. Since mid-June, speculative articles about Qualcomm acquiring AI chip startup Tensorrent for $8 billion to $10 billion had been circulating in the industry, but Tensorrent CEO Jim Keller himself explicitly stated, "There have been no such negotiations." Meanwhile, at the same time, Qualcomm announced a definite acquisition of another company. This article will examine this "gap between rumor and reality."

The Origin: A Mid-June Speculative Article

The story began with an article published on June 15th by industry media outlet "The Information." The article claimed that Qualcomm was in talks to acquire Canadian AI chip startup Tensorrent for $8 billion to $10 billion. Reuters followed suit, and the market took the news quite seriously, with Qualcomm shares rising over 4% in after-hours trading at one point.

Tensorrent itself is a well-known name in the industry. CEO Jim Keller is a legendary engineer in the chip industry, having worked on AMD's K8 architecture (the design that formed the basis of the Athlon 64), Apple's A4 and A5 processors (used in the original iPad and iPhone 4), AMD's Zen architecture (a design said to have ended Intel's dominance), and even Tesla's self-driving car computer. The company develops AI accelerators based on the open instruction set architecture "RISC-V," and has an open-source software platform called "TT-Metalium" to compete with closed software stacks like NVIDIA's CUDA.

What awaited at the investor event on June 24th was "another announcement"

After the speculative articles, many industry watchers expected an official announcement at Qualcomm Investor Day on June 24th. However, upon closer inspection, there was no official announcement regarding Tensorrent. Instead, Qualcomm announced the acquisition of Modular, an AI inference software startup.

Modular develops the "Mojo" programming language and the "MAX" inference engine, which allow AI models to run on chips from different manufacturers such as NVIDIA, AMD, Intel, and Qualcomm without rewriting code. The acquisition price was approximately $3.92 billion (all stock exchange), and this transaction has been officially confirmed. At the same investor event, Qualcomm also revealed plans to begin shipping custom silicon to major cloud providers by 2026.

While reports continued to indicate that Tensorrent acquisition negotiations were "ongoing," the subsequent developments are intriguing.

And Now, Keller Himself Denies Negotiations

According to recent reports, Tensorrent CEO Jim Keller himself has stated that no acquisition negotiations are taking place with Qualcomm, denying previous reports of an $8 billion to $10 billion deal.

From the initial speculative articles, the reports consistently included the caveat that "discussions are ongoing but nothing has been finalized." Keller's denial now marks a clear conclusion, with the rumors being debunked by the person himself. Of course, it's unclear whether any behind-the-scenes discussions never existed at all, or whether they were temporary but have since fizzled out. However, it's important to note that, at the very least, this is not an acquisition that will be officially announced anytime soon.

Why Did This Rumor Gain So Much Attention?

Understanding the technical background helps explain why this rumor was taken so seriously.

NVIDIA GPUs are overwhelmingly powerful in applications that require simultaneous, large-scale parallel processing, such as large-scale training. However, in actual "inference"—responding to user requests—especially for small-scale processing such as handling one request at a time, the parallel processing capabilities of GPUs cannot be fully utilized, leading to a weakness in power efficiency. Tenstorrent's chips have a design philosophy that addresses this weakness in inference-oriented processing, and have attracted attention as a potential challenger to NVIDIA's dominance.

Bernstein analysts cited "talent retention" as the biggest risk if the acquisition were to materialize. Keller has a history of leaving companies after only a few years, including AMD, Apple, and Intel, and it's unlikely he would stay long even if the acquisition went through. Qualcomm's previous acquisition of Ventana Micro Systems, which also employed Keller, saw many engineers leave after the lock-up period ended. While the negotiations themselves have been denied by Keller, this "talent risk" is likely to remain a recurring issue even if some deal were to emerge in the future.

Things Engineers Should Notice

What I personally find interesting about this whole sequence of events is that the "confirmed fact" of the Modular acquisition and the "denied rumor" of the Tensorrent acquisition were reported almost simultaneously. Silicon Valley corporate acquisition reports often have a trial balloon aspect (information is deliberately released to gauge market reaction). Whether this is the case here is unclear, but I think it served as a reminder of the importance of clearly distinguishing between confirmed information and speculative articles—a basic but often forgotten attitude.

Qualcomm's inference chip strategy itself is steadily progressing in the form of the Modular acquisition. Whether this will lead to a full-scale entry into the RISC-V-based open hardware field, or whether this rumor will end there, is something I'd like to continue following Qualcomm's movements.

QualcommTenstorrentModularRISC-VAIチップJim Keller
Advertisement300 × 250