Friday, October 9, 2026 Trend Press · Cloudflare Pages

The Trend Tribune

"All the trends that are fit to read" Evening Edition Free of Charge
TODAY'S LEAD STORY

GitHub with 722 "manuscripts"—OpenAI demonstrates the practice of public access that lends credibility to AI mathematics.

On October 6th, OpenAI released 722 mathematical manuscripts (372 families) based on its previously unpublished in-house frontier model on GitHub. This article will analyze the revision and citation protocols aligned with the recommendations of the Institute for Advanced Study's advisory group, the scope and limitations of the Lean format, the denominator of the trial (approximately 4,000 problems), the constraint of non-reproducibility due to the unpublished model, and OpenAI's commitment to self-improvement in quality.

GitHub with 722 "manuscripts"—OpenAI demonstrates the practice of public access that lends credibility to AI mathematics.
(Photo: illustrative)

722 "Manuscripts" Released on GitHub

On October 6th, OpenAI released a research paper titled "Sharing AI progress in mathematics," along with a collection of mathematical results generated by its previously unreleased in-house frontier model. The files are located in the "openai/math" repository on GitHub under the Apache-2.0 license. The collection consists of 722 manuscripts, organized into 372 related "families." The topics covered are broad, including number theory, computational complexity, geometry, and mathematical physics. While the magnitude of the mathematical achievements is significant, what caught my attention as a researcher was the "method" of presenting the results. How do you gain the trust of the mathematical community for AI-generated mathematics? This collection reveals OpenAI's exploration of this process, often with external advice.

Advice from the Institute for Advanced Study Group

OpenAI explains that it consulted with an independent "Advisory Group on Mathematics and Artificial Intelligence" at the Institute for Advanced Study in Princeton regarding this publication method, and referred to their previously published recommendations. There is already a movement from mathematicians to standardize the publication practices for AI-generated results, and OpenAI designed its publication method in line with this. Specifically, it has established protocols for manuscript revision and citation, and ensures that older versions remain accessible. Regarding the paper's hosting, in addition to GitHub, they are exploring community-based hosting options that align with the advisory group's guidelines.

The Meaning and Limitations of "Lean Verified"

The Lean-based formalization is prominently featured as a guarantee of reliability. Lean is a programming language that allows mathematical proofs to be mechanically checked by computers. A formalized proof will fail from the start if there are leaps or holes in the logical flow. OpenAI has published Lean versions of many proofs and plans to add more as they become available. However, it's important to note that this is "many," not "all." Furthermore, as a general principle, Lean guarantees that "formalized propositions have been proven," but whether those propositions precisely match the theorems intended by humans needs to be verified separately by humans. Mistakes at the formalization stage cannot be detected by machine inspection alone.

Seeing the "Contents" of Transparency in Numbers

OpenAI also provides information on the process leading to the results. This includes 10 documents summarizing the model's inference, an estimate of the computational cost converted to the use of ChatGPT Pro, and statistics on the number of problems attempted. According to the report, approximately 4,000 problems were involved in the evaluation, and the average computational cost per result is equivalent to about 3 hours of thinking on ChatGPT Pro. The README also states that the evaluation was extended to unsolved research problems because existing mathematical evaluations had reached saturation. What we need to calmly consider here is that we cannot simply divide approximately 4,000 problems by 372 families to obtain a "success rate." This is because one problem may lead to multiple documents, and vice versa. While the disclosure of the denominator of the trials is commendable, caution is needed in interpreting it.

The Biggest Constraint: Inability to Reproduce

On the other hand, the biggest constraint from a verification perspective is clear. The model that produced the results remains undisclosed, making it impossible for external mathematicians to reproduce and verify the same procedure. OpenAI itself only states that it is working to release this model in a responsible manner. In other words, at present, anyone can verify the parts that have been machine-checked with Lean, but it is impossible to verify from the outside "how the proof was arrived at" or "whether similar results can be obtained for other problems under the same conditions."

Self-Identified Manuscript Quality as a Challenge

Another point that reveals their frankness is the list of areas for future improvement. OpenAI has promised to improve the quality of citations, the way mathematical explanations are presented, and the way results are presented, in preparation for future publication. Conversely, this means that there is room for improvement in these areas of the current manuscript. A research paper is not sufficient if the proof is correct; it can only be used as community knowledge when it includes context within previous research and explanations that readers can understand. Given the sheer volume of 722 papers, whether the quality aligns with the volume remains to be seen, and we must await individual evaluations by mathematicians.

What Researchers Should Keep an Eye On

Ultimately, the value of the results will be determined by how mathematicians in each field evaluate each paper. OpenAI has also indicated plans to fund workshops, conferences, and special programs focused on understanding the major achievements of AI. What's important about this release is not the sheer size of the results themselves, but the fact that a "framework" has begun to emerge for presenting AI-generated mathematics in a way that the community can verify, cite, and critique. We will continue to closely monitor whether the Lean formatting is expanded, models are made available externally, and verification results from independent mathematicians become readily available.

OpenAI数学LeanAI for Science企業公式発表

What Vibration-Dampening Floors and Large Cranes Tell Us: A Read on the Full Opening of the GMO Humanoid Lab Without the "Company Research" Label

On October 6th, GMO Internet Group fully opened its "GMO Humanoid Lab" in Shibuya. We will evaluate the development area equipped with vibration-damping floors and large cranes, the division of roles among the three companies, and the trading company-type strategy of using overseas-made robots (such as Unitree) and focusing on peripheral services (security evaluation, data return, and maintenance vehicles), while also verifying that the claims of being "one of the largest in Japan" and "the first in Japan" are both based on the company's own research.

Vibration-Dampening Floors and Large Cranes: Equipment That Will Impress Robot Engineers

On October 6th, GMO Internet Group fully opened the "GMO Humanoid Lab Shibuya Showroom" on the 11th floor of the Cerulean Tower in Shibuya, Tokyo. The facility, which had opened on April 7th with approximately half of its 382 tsubo (approximately 1,260 square meters) floor space, is now fully operational with the addition of the remaining area. What caught the eye most in the press release wasn't the glamorous interior design, but the statement that the newly established development area includes "vibration-dampening floors" and "large cranes." Anyone with experience in humanoid robot development will immediately understand the purpose of these two pieces of equipment.

Why Floors and Cranes?

Experiments with humanoid robots are inherently risky. During walking control and whole-body control tests, the robot can tip over every time the control law is changed. Therefore, during testing, it's common practice to suspend the robot with a crane to protect the expensive robot and personnel in the event of a fall. The vibration-damping floor is likely designed to prevent vibrations from heavy robots walking, jumping, and experiencing impacts from affecting work or measurements in adjacent areas. While this includes some speculation on my part, both facilities seem more like "facilities for repeated experimentation" than mere "exhibition equipment," clearly demonstrating a genuine intention to make this base a place for serious development, not just exhibitions.

Division of Roles Among the Three Companies

At the base, three companies—GMO Internet Group, GMO AI & Robotics Trading (GMO AIR), and GMO Various Robotics—work together. The Group's Research and Development Headquarters is responsible for exploring and verifying new technologies through the research and implementation of cutting-edge research papers. GMO AIR handles everything from selecting and procuring advanced robotic units domestically and internationally to planning implementation projects and developing use cases. GMO Various Robotics is responsible for research and development of humanoid and physical AI, as well as the development and demonstration of solutions using robot control. The company, founded in January 2025, claims to have achieved autonomous driving at a top speed of 250 km/h in the autonomous driving formula race "A2RL" and won the Silver Race in its first appearance. Up to this point, these are self-introductions from each company and not figures verified by a third party, but the division of roles into aircraft procurement, research, and demonstration makes sense structurally.

The Content of "Japan's Largest"

The press release calls this facility "Japan's largest physical AI research and development base." The accompanying note is interesting. The basis is stated as "GMO Internet Group survey. As of October 6, 2026, domestic web survey." In other words, it's not a certification by a third party, but a statement based on their own web survey. While the figure of 382 tsubo (approximately 1,250 square meters) is an easily verifiable fact, whether to take the "largest" label at face value is another matter, and readers should take a step back and consider it carefully. At the soft opening in April, it was reported that they aimed for "100 robots and 100 engineers." This release doesn't mention the current number of units or personnel. It's important to distinguish between goals and the current situation.

Overseas-made robots, the focus is on "peripheral" areas

The robots gathered at the base are advanced robots, including humanoids from Unitree Robotics. In other words, their strategy isn't to manufacture the hardware itself. Instead, GMO is focusing on the peripheral areas that support the robots after deployment. This includes "GMO SAFE for Humanity," which assesses security risks according to the model and usage environment, and provides support for monitoring during operation; "GMO LOOP for Physical AI," which uses data obtained in the field to enhance models and improve actual robots; and the "GMO Humanoid Ambulance," where engineers arrive in a dedicated vehicle in case of malfunction or trouble to diagnose, repair, and provide replacement units. The ambulance is said to be the first of its kind in Japan, although this "first" claim is also based on the company's own research. It's a trading company model, bringing its strengths in internet infrastructure and security to the operation side of the humanoids. Considering the domestic situation where hardware manufacturing is not feasible, this seems like a realistic approach.

What We Want to Verify: "Is It Working?"

As a robotics engineer, what I want to confirm is whether these services are actually running. The claim that LOOP continuously improves operational accuracy can only be evaluated if the improvement is shown in numbers. Similarly, with ambulances, the value of the maintenance service becomes clear once results such as the number of dispatches and recovery time are available. Regarding SAFE, if the evaluation criteria and which robots it was applied to are specifically disclosed, it would provide useful guidance for companies considering implementation.

What Engineers Should Look For

GMO has positioned 2026 as the "Year One of Humanoids" and explains that the base is a place to move from "research" to "social implementation." Judging from the contents of the equipment, their commitment to repeated experimentation seems genuine. On the other hand, the basis for claims such as "largest" and "first in Japan" is limited to their own research. Next, I want to focus on what kind of actual implementation cases and publicly available data will emerge from this base. I want to evaluate it based on the numbers accumulated on-site, not just the claims.

GMOヒューマノイドフィジカルAIUnitree国産ロボット

The day when "selectable models" become one: The meaning of Gemini's free tier becoming Flash-Lite only

Starting October 9th, Google will limit the available models for free users of its personal Gemini service to only the lightweight Flash-Lite version. Flash and Pro will be removed from the free tier, the $4.99/month AI Plus plan will lose its Pro version, and Deep Think will be moved to the two higher-tier plans. This article will examine the meaning of this design, which uses "selectable models" rather than rate limits, and the lack of reports regarding the impact on the API.

The Day When Only One Model is Available

On October 9th, Google will significantly narrow its free Gemini tier for personal use. This change, reported by 9to5Google on October 3rd and subsequently followed by multiple media outlets, will limit free users to only one model: the lightweight "Flash-Lite" (some reports say Gemini 3.5 Flash-Lite). "Gemini 3.6 Flash," which was previously available even for free users, and "Gemini 3.1 Pro," which had limited usage, will disappear from the model picker. Instead of tightening rate limits, the focus is on drawing a line by changing the "types of models available." This design shift is the essence of this change.

A New Map for Each Plan

Comparing the reports, the new configuration is roughly as follows: Free users will only have Flash-Lite. The cheapest paid plan, "Google AI Plus," reportedly at $4.99 per month, will offer a choice between Flash-Lite and Flash, but the previously available Pro version will be removed. To use Pro, you'll need "Google AI Pro" (reportedly $19.99/month) or higher, and "Deep Think," which allows for more complex reasoning, will also be moved to the top two plans: AI Pro and Ultra. AI Plus users will receive an email notifying them when the changes will be reflected in their accounts. This applies only to individual accounts.

"Effort Level" No Loophole

Technically interesting is the relationship with the "Effort" setting, which adjusts the depth of thinking. According to Nokiamob's explanation, even if a free user selects high effort with Flash-Lite, they will still be using Flash-Lite and won't get behavior close to Pro. This is natural, given the difference in model size. However, for those who previously could try Pro several times with the free version, the ability to "rely on higher-end models only when asking difficult questions" will no longer be available. Heavy tasks that Pro is intended for, such as mathematics, coding, and multi-document analysis, have clearly been moved to the paid side.

Untagged Costs Appear in the Menu

Why now? No official explanation has been reported, and the only prominent explanations are that it's a "strategic move to divide capabilities into tiers and encourage upgrades." However, there are some clues to infer the background. The computational cost of inference becomes a burden on the free tier as the number of users increases. Moreover, another report states that major memory manufacturers have shifted their supply to AI data centers, resulting in a roughly 60% drop in shipments of smartphones priced under $100. Computational resources and the components that support them are becoming "shortages" globally. In this situation, it is likely that continuing to offer Pro and Flash for free is becoming unsustainable from a business perspective.

The convergence of the "lightweight model for free" industry

Broadening the perspective, this doesn't seem to be a problem limited to Google alone. In the AI ​​industry, a configuration is spreading where lightweight, low-cost models are assigned to the free version, and higher-end models to the paid version. If the intelligence of the models available for free and paid versions differs, it is a natural consequence that the quality of answers and the frequency of factual errors will differ even for the same question. In the future, we will likely see even more instances where the reputation of AI performance is divided depending on "which plan was tested."

Impact on Developers: Unclear from News Reports

On the other hand, there are still some unclear points that need clarification. The recent reports all focus on model selection in the Gemini app and web version, and do not touch upon the free tier via API or its handling in Google AI Studio. It's currently impossible to determine whether this change will directly affect developers prototyping via API. For the time being, it's safest to check for updates to the official documentation.

Things Engineers Should Keep an Eye On

The biggest change is that the "number of usable models" has become the differentiating factor for pricing plans. Evaluating Gemini's capabilities solely based on the free version will lead to misjudging the difference between it and the paid Pro version. If non-engineers around you say, "Is this all the AI ​​can do?", it's a good idea to ask them which plan and which model they used. We should continue to closely monitor how the experience of free users changes after October 9th, how much migration to paid plans progresses, and how competitors modify the model configurations in their free versions.

GoogleGemini無料プラン料金プランLLM
Advertisement300 × 250