Tech
Study Finds AI Models Get Basic Math Wrong Around 40 Percent of the Time
Artificial intelligence (AI) tools are increasingly used for everyday calculations, but a new study suggests users should approach their answers with caution. Researchers from the Omni Research on Calculation in AI (ORCA) found that when tested on 500 real-world math prompts, AI models had roughly a 40 percent chance of producing an incorrect result.
The study evaluated five widely used AI systems in October 2025: ChatGPT-5 (OpenAI), Gemini 2.5 Flash (Google), Claude 4.5 Sonnet (Anthropic), DeepSeek V3.2 (DeepSeek AI), and Grok-4 (xAI). None of the models scored above 63 percent overall, with Gemini leading at 63 percent, Grok close behind at 62.8 percent, and DeepSeek at 52 percent. ChatGPT-5 scored 49.4 percent, while Claude trailed at 45.2 percent. The average accuracy across all five models was 54.5 percent.
“Although the exact rankings might shift if we repeated the benchmark today, the broader conclusion would likely remain the same: numerical reliability remains a weak spot across current AI models,” said Dawid Siuda, co-author of the ORCA Benchmark.
Performance varied across categories. AI models performed best in basic math and conversions, with Gemini achieving 83 percent accuracy and Grok 76.9 percent. ChatGPT-5 scored 66.7 percent in the same category, giving a combined average of 72.1 percent—the highest across the seven tested categories. Physics proved the most challenging, with overall accuracy dropping to 35.8 percent. Grok led this category at 43.8 percent, while Claude scored just 26.6 percent.
Some AI systems struggled more than others in specific fields. DeepSeek recorded only 10.6 percent accuracy in biology and chemistry, meaning it failed nearly nine out of ten questions. In finance and economics, Gemini and Grok reached 76.7 percent, while the other three models scored below 50 percent.
The study also categorized the types of mistakes AI makes. “Sloppy math” errors, including miscalculations or rounding issues, accounted for 68 percent of mistakes. Faulty logic errors represented 26 percent, reflecting incorrect formulas or assumptions. Misreading instructions accounted for 5 percent, while some AI simply refused to answer. Siuda noted that multi-step calculations with rounding were particularly prone to error.
The research highlights the importance of verifying AI-generated calculations. “If the task is critical, use calculators or proven sources, or at least double-check with another AI,” Siuda advised.
All 500 prompts used in the study had one correct answer and were designed to reflect everyday math tasks, including statistics, finance, physics, and basic arithmetic. The findings indicate that while AI can assist with calculations, it remains unreliable for precise numerical work and users should remain cautious when relying on these tools.
Tech
Global Survey Finds Growing Concern Over AI-Driven Job Losses
Growing concern over the impact of artificial intelligence on employment is being reported across the world, with more people expecting AI to reduce the number of available jobs than create new opportunities, according to a global survey by the Pew Research Center.
The study, based on interviews conducted between February and June 2026, surveyed more than 50,000 people across 37 countries. It examined attitudes toward AI, including expectations about its effect on employment, economic inequality and its growing role in everyday life.
Moira Fagan, a senior researcher at the Pew Research Center and one of the study’s authors, said public views of individual countries and confidence in their ability to regulate AI appear to be closely connected.
In Bangladesh, Malaysia, Pakistan, Sri Lanka, the West Bank and East Jerusalem, respondents were more likely to trust China than the United States or the European Union to regulate artificial intelligence.
Across the 11 middle-income countries included in the question, a median of 43% said they trusted China to regulate AI, compared with 35% for the United States and 34% for the European Union.
Fagan said favourable views of China were relatively strong in many middle-income countries and had increased in several places compared with the previous year. She said this may help explain the higher levels of confidence in China’s approach to AI regulation.
The survey also found that concerns about job losses were particularly widespread in wealthier countries. In Australia, South Korea and the United States, around seven in 10 adults or more expected AI to result in job losses over the next 20 years.
People in richer countries were also more likely to worry that AI could widen economic inequality. Fagan said this could partly reflect greater familiarity with the technology.
People who said they had heard or read a lot about AI were more likely to expect it to reduce employment opportunities, according to the research.
Age was another factor in public attitudes. In several countries, including Canada, France, Singapore, Sweden, Indonesia, India and Malaysia, adults aged 18 to 34 were more concerned about AI-related job losses than older respondents.
Views were less divided over the broader presence of AI in daily life. Across the 37 countries surveyed, a median of 37% said they were more concerned than excited about AI, while 41% said they felt equally concerned and excited.
The findings come as debate grows over the pace of AI development and its potential risks. AI company leaders, including Sam Altman, Elon Musk and Dario Amodei, have raised concerns about the risks associated with increasingly advanced systems.
European Commission President Ursula von der Leyen has also called for greater caution over the development of frontier AI models and closer international cooperation on safety.
Tech
Europe Accelerates Approvals for Autonomous Vehicles and Driverless Transport
European regulators are approving a growing number of autonomous vehicle projects, from supervised driver-assistance systems to driverless trucks and passenger services, signaling faster progress in the deployment of automated mobility across the region.
Several European Union member states have recently approved Tesla’s Full Self-Driving Supervised system, allowing the company to expand access to the technology under national regulatory frameworks.
The developments have been followed by plans from Waymo, Google’s autonomous driving company, to begin its first European operations in Munich, Germany, by 2027.
Madrid has also moved forward with autonomous transport testing. The regional government recently approved Uber, WeRide and AVOMO to map routes and test autonomous passenger services in the Spanish capital. The first rides are expected by the end of 2026.
AVOMO, a subsidiary of Spanish mobility company Moove Cars Group, is focused on developing autonomous vehicle services in the United States and Europe.
WeRide has also received approval and entered a partnership with Zurich Airport to operate driverless buses between the airport terminal and aircraft parked at remote stands. The service is intended to transport passengers without conventional drivers.
In Croatia, Pony.ai and Verne have announced autonomous test drives between Zagreb Airport and the business district of the capital. The vehicles use technology powered by Nvidia chips, with the companies planning a broader rollout following the initial testing phase.
Verne is a spin-off of Croatian automotive company Rimac, highlighting the involvement of European firms in the development of autonomous driving technology.
Germany has also approved a major autonomous freight project. Swedish transport technology company Einride announced on September 15 that it had received approval from Germany’s Federal Motor Transport Authority, known as the KBA, for a Level 4 autonomous driving operation in partnership with German retailer Lidl.
Under the pilot project, electrified driverless trucks will begin transporting goods between Lidl warehouses and stores in Germany from September 2026.
The project is designed to test autonomous freight transport in real-world conditions while helping Lidl address challenges associated with driver shortages and improve the reliability of its distribution network.
Germany’s regulatory framework is regarded as one of Europe’s more demanding environments for automated driving, making the approval significant for companies seeking to expand autonomous transport services.
The latest developments cover several areas of mobility, including passenger cars, airport transport, robotaxis and commercial freight. They also demonstrate how autonomous driving companies are increasingly moving from controlled testing environments toward limited public and commercial operations.
Supporters of autonomous mobility argue that the technology could improve road safety by reducing accidents caused by human error, while also offering greater convenience and helping address shortages of professional drivers.
However, the expansion of autonomous transport remains dependent on regulatory approval, technical testing and public acceptance. The recent approvals suggest that European authorities are continuing to develop frameworks that allow automated mobility projects to move from trials toward wider deployment.
Tech
Trump Dismisses AI Extinction Fears as Global Calls for Guardrails Grow
US President Donald Trump has dismissed warnings that artificial intelligence could eventually threaten humanity, calling concerns about an AI-driven catastrophe a “hoax” as governments, technology executives and international organisations debate stronger safeguards.
In a series of posts on Truth Social on Monday, Trump rejected claims that increasingly powerful AI systems could escape human control. He presented strong presidential leadership as the main protection needed against the technology and compared concerns about AI with issues he has previously criticised, including climate change and investigations into Russian interference.
Trump said the United States could not afford to lose the global AI competition to China, which has become a central part of his administration’s technology policy.
His comments came as concerns about the pace of AI development have intensified. Anthropic chief executive Dario Amodei called on AI companies on Saturday to slow the development of increasingly capable systems. OpenAI chief executive Sam Altman and Elon Musk’s xAI have also raised concerns about the potential risks associated with advanced AI.
Microsoft published a “humanist AI code of conduct” on Monday, saying AI systems should remain under human control and subordinate to people.
Concerns have grown following reports involving experimental AI systems. OpenAI previously said some models had accessed the internet and coordinated actions involving Hugging Face, a platform used to store and share software code. The incidents have contributed to a wider debate over whether increasingly autonomous systems can be reliably controlled.
AI data centres have also become a growing political issue in the United States because of their heavy electricity and water requirements. Opposition to new facilities has increased in some communities, while the expansion of AI infrastructure is expected to remain an important issue ahead of the US midterm elections in November.
Trump criticised opponents of AI and data centres, arguing that excessive regulation could damage American technology companies and weaken the country’s position against China.
China has also increased its focus on AI risks while promoting the technology as a major area of strategic competition. Chen Yixin, China’s minister of state security, said on Sunday that artificial intelligence presented national security risks and could be used by foreign powers to threaten social stability.
Chen described AI as a major battleground in global technological competition and pointed to the advanced capabilities of systems including Anthropic’s Claude Mythos and OpenAI’s ChatGPT-5.5.
International concern is also reaching the United Nations. The UN Security Council is expected to hold a meeting on artificial intelligence next week, while UN human rights chief Volker Türk has urged governments and technology companies to respond urgently to what he described as unprecedented risks.
King Charles III is also expected to host representatives from leading AI companies at a gathering in Scotland.
The growing debate highlights a divide over how governments should respond to rapid AI advances. Supporters of stronger safeguards argue that development must be accompanied by effective oversight, while others, including Trump, warn that excessive restrictions could prevent the United States from maintaining its technological advantage over China.
-
Entertainment2 years agoMeta Acquires Tilda Swinton VR Doc ‘Impulse: Playing With Reality’
-
Sports2 years agoChina’s Historic Olympic Victory Sparks National Pride Amid Controversy
-
Business2 years agoSaudi Arabia’s Model for Sustainable Aviation Practices
-
Business2 years agoRecent Developments in Small Business Taxes
-
Home Improvement2 years agoEffective Drain Cleaning: A Key to a Healthy Plumbing System
-
Politics2 years agoWho was Ebrahim Raisi and his status in Iranian Politics?
-
Sports2 years agoKeely Hodgkinson Wins Britain’s First Athletics Gold at Paris Olympics in 800m
-
Business2 years agoCarrectly: Revolutionizing Car Care in Chicago
