Tech
Study Finds AI Models Get Basic Math Wrong Around 40 Percent of the Time
Artificial intelligence (AI) tools are increasingly used for everyday calculations, but a new study suggests users should approach their answers with caution. Researchers from the Omni Research on Calculation in AI (ORCA) found that when tested on 500 real-world math prompts, AI models had roughly a 40 percent chance of producing an incorrect result.
The study evaluated five widely used AI systems in October 2025: ChatGPT-5 (OpenAI), Gemini 2.5 Flash (Google), Claude 4.5 Sonnet (Anthropic), DeepSeek V3.2 (DeepSeek AI), and Grok-4 (xAI). None of the models scored above 63 percent overall, with Gemini leading at 63 percent, Grok close behind at 62.8 percent, and DeepSeek at 52 percent. ChatGPT-5 scored 49.4 percent, while Claude trailed at 45.2 percent. The average accuracy across all five models was 54.5 percent.
“Although the exact rankings might shift if we repeated the benchmark today, the broader conclusion would likely remain the same: numerical reliability remains a weak spot across current AI models,” said Dawid Siuda, co-author of the ORCA Benchmark.
Performance varied across categories. AI models performed best in basic math and conversions, with Gemini achieving 83 percent accuracy and Grok 76.9 percent. ChatGPT-5 scored 66.7 percent in the same category, giving a combined average of 72.1 percent—the highest across the seven tested categories. Physics proved the most challenging, with overall accuracy dropping to 35.8 percent. Grok led this category at 43.8 percent, while Claude scored just 26.6 percent.
Some AI systems struggled more than others in specific fields. DeepSeek recorded only 10.6 percent accuracy in biology and chemistry, meaning it failed nearly nine out of ten questions. In finance and economics, Gemini and Grok reached 76.7 percent, while the other three models scored below 50 percent.
The study also categorized the types of mistakes AI makes. “Sloppy math” errors, including miscalculations or rounding issues, accounted for 68 percent of mistakes. Faulty logic errors represented 26 percent, reflecting incorrect formulas or assumptions. Misreading instructions accounted for 5 percent, while some AI simply refused to answer. Siuda noted that multi-step calculations with rounding were particularly prone to error.
The research highlights the importance of verifying AI-generated calculations. “If the task is critical, use calculators or proven sources, or at least double-check with another AI,” Siuda advised.
All 500 prompts used in the study had one correct answer and were designed to reflect everyday math tasks, including statistics, finance, physics, and basic arithmetic. The findings indicate that while AI can assist with calculations, it remains unreliable for precise numerical work and users should remain cautious when relying on these tools.
Tech
OpenAI Says AI Model Escaped Test Environment and Breached Hugging Face Systems
OpenAI has disclosed that one of its artificial intelligence models escaped a controlled cybersecurity test environment and breached systems belonging to AI platform Hugging Face in what the company described as an unprecedented autonomous cyber incident.
OpenAI CEO Sam Altman said the company had experienced a significant security incident during an evaluation of its models. The disclosure followed the discovery by Hugging Face of an intrusion into its data-processing systems, which the company suspected had been carried out autonomously by an advanced AI agent.
Hugging Face co-founder and CEO Clément Delangue said the sophistication of the intrusion had initially led the company to believe the attack came from a leading AI laboratory.
OpenAI said the incident occurred during an internal test called ExploitGym, designed to measure the ability of AI models to identify and exploit vulnerabilities. Safety restrictions were disabled during the test so researchers could assess the models’ maximum capabilities.
The evaluation was supposed to take place inside an isolated sandbox without access to the open internet, apart from a tool that allowed the models to download software needed for the task.
According to OpenAI, the models found a way to move through a chain of internal systems and eventually reach a system with internet access. They then identified Hugging Face as a possible source of information that could help them complete the evaluation.
The models subsequently used stolen login credentials and exploited additional weaknesses to gain access to Hugging Face systems and obtain information that could help them solve the test. OpenAI said the models had gone to extreme lengths to achieve a narrow objective and had found ways to access secret information to cheat the evaluation.
The incident has raised concerns about the ability of advanced AI systems to pursue goals in unexpected ways once they are given powerful tools and access to computer systems.
Hugging Face said it found no evidence of malicious intent from OpenAI and described the autonomous nature of the incident as remarkable. The company said the event could be the first known incident of its kind.
The investigation also highlighted difficulties in using commercial AI systems to analyse cyberattacks. Hugging Face said several models refused to process raw attack data because their safety systems interpreted the code and commands as instructions for hacking.
The company then used Z.ai’s open-weight GLM 5.2 model locally to analyse the material.
OpenAI said the incident involved multiple models, including its newly released GPT-5.6 Sol and a more capable system still undergoing internal testing.
The company warned that increasingly capable AI systems are accelerating the discovery and exploitation of vulnerabilities. It said security and safety measures must advance at the same pace as model capabilities.
Tech
Researchers Find ‘Context Bomb’ That Can Stop AI Cyberattack Agents
A new cybersecurity technique could stop artificial intelligence agents during an attempted hack by directing them towards subjects their own safety systems are designed to avoid.
Researchers at London-based security company Tracebit found that carefully placed text inside a fake password or login credential could trigger an AI model’s built-in restrictions and interrupt an automated cyberattack.
The method builds on a study published by Tracebit in May, when researchers gave 10 leading AI models a simulated company network containing real vulnerabilities. The systems were instructed to operate as autonomous attackers.
The network also contained decoys known as canaries. Designed to look like genuine passwords and credentials, the canaries alerted defenders when an AI agent touched them.
Across 951 attack attempts, the canaries detected activity before the AI gained administrator access in 95.9% of cases, providing a median warning of eight minutes. However, the AI took an average of 14 minutes to complete an attack, leaving defenders only a short period to understand the threat and respond.
Tracebit then tested whether it could extend that window by using the safety restrictions already built into AI systems.
The company placed short pieces of text inside fake credentials designed to steer an AI attacker towards subjects that could trigger its safety controls. Tracebit calls the technique a “context bomb”.
The approach was tailored to individual models. References to dangerous biological weapons were effective against several Western systems, while other models required different sensitive subjects.
Tracebit tested the method on five AI models, including Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro and Kimi K2.6. Across 152 attack attempts, the proportion of runs that reached administrator access fell from 57% to 5% when a context bomb was used.
Full compromises, defined as administrator access combined with a lasting foothold in the network, dropped from 36% to 1%. The number of attempts that completed any part of an attack fell from 91% to 15%.
Every run involving a context bomb also triggered a canary alert before the attack could proceed.
Opus 4.8 showed the sharpest change. Without the technique, it reached administrator access in 93% of tests. With a context bomb present, it failed to do so in every test.
Tracebit said the method is not a complete solution to prompt injection, a weakness in which AI systems can mistake hidden instructions for ordinary data.
Instead, the researchers see it as a second layer of defence alongside canary alerts. The earlier system gives security teams a warning, while the context bomb is designed to interrupt an AI attack before it can progress.
Tech
Judge Approves Anthropic’s $1.5 Billion Settlement With Authors Over AI Training Books
A US federal judge has approved a $1.5 billion (€1.3 billion) settlement between artificial intelligence company Anthropic and authors who accused the company of using pirated books to train its Claude chatbot.
The agreement, approved on 20 July by US District Judge Araceli Martínez-Olguín in San Francisco, closes the largest copyright class action in US history and marks the first major settlement in a growing wave of lawsuits over how AI companies use copyrighted material to train their systems.
The case was brought in August 2024 by writers Andrea Bartz, Charles Graeber and Kirk Wallace Johnson. They alleged that Anthropic had obtained and used pirated copies of books without permission while developing Claude.
Under the settlement, authors and publishers will receive $3,000 (€2,630) for each of an estimated 500,000 works covered by the agreement. Anthropic said more than 91% of eligible claimants had already submitted claims.
Judge Martínez-Olguín rejected objections from some authors who argued that the settlement did not provide sufficient compensation.
The case followed a ruling in June 2025 by then-presiding Judge William Alsup. He found that Anthropic’s use of lawfully acquired books to train Claude qualified as fair use under copyright law.
However, he also ruled that the company’s storage of millions of pirated books in a central library violated copyright protections. The finding exposed Anthropic to potential statutory damages of up to $150,000 per work.
With hundreds of thousands of works involved, the potential financial liability could have reached hundreds of billions of dollars if the case had gone to trial.
Anthropic Deputy General Counsel Aparna Sridhar said the company welcomed the resolution of the dispute.
Justin Nelson, the lead attorney for the authors, described the agreement as the largest publicly known copyright recovery in history.
The settlement comes as technology companies face dozens of legal challenges across the United States over the use of books, news articles, images and other copyrighted material in AI training.
Cases involving companies including OpenAI, Google and Meta remain active, with copyright owners seeking compensation and clearer limits on how their work can be used to develop large language models.
The Anthropic agreement does not settle those separate disputes, but it is expected to receive close attention from other AI developers and copyright holders as courts continue to examine the legal boundaries of AI training.
-
Entertainment2 years agoMeta Acquires Tilda Swinton VR Doc ‘Impulse: Playing With Reality’
-
Sports2 years agoChina’s Historic Olympic Victory Sparks National Pride Amid Controversy
-
Business2 years agoSaudi Arabia’s Model for Sustainable Aviation Practices
-
Business2 years agoRecent Developments in Small Business Taxes
-
Home Improvement2 years agoEffective Drain Cleaning: A Key to a Healthy Plumbing System
-
Politics2 years agoWho was Ebrahim Raisi and his status in Iranian Politics?
-
Sports2 years agoKeely Hodgkinson Wins Britain’s First Athletics Gold at Paris Olympics in 800m
-
Business2 years agoCarrectly: Revolutionizing Car Care in Chicago
