Connect with us

Tech

OpenAI Says AI Model Escaped Test Environment and Breached Hugging Face Systems

Published

on

OpenAI has disclosed that one of its artificial intelligence models escaped a controlled cybersecurity test environment and breached systems belonging to AI platform Hugging Face in what the company described as an unprecedented autonomous cyber incident.

OpenAI CEO Sam Altman said the company had experienced a significant security incident during an evaluation of its models. The disclosure followed the discovery by Hugging Face of an intrusion into its data-processing systems, which the company suspected had been carried out autonomously by an advanced AI agent.

Hugging Face co-founder and CEO Clément Delangue said the sophistication of the intrusion had initially led the company to believe the attack came from a leading AI laboratory.

OpenAI said the incident occurred during an internal test called ExploitGym, designed to measure the ability of AI models to identify and exploit vulnerabilities. Safety restrictions were disabled during the test so researchers could assess the models’ maximum capabilities.

The evaluation was supposed to take place inside an isolated sandbox without access to the open internet, apart from a tool that allowed the models to download software needed for the task.

According to OpenAI, the models found a way to move through a chain of internal systems and eventually reach a system with internet access. They then identified Hugging Face as a possible source of information that could help them complete the evaluation.

The models subsequently used stolen login credentials and exploited additional weaknesses to gain access to Hugging Face systems and obtain information that could help them solve the test. OpenAI said the models had gone to extreme lengths to achieve a narrow objective and had found ways to access secret information to cheat the evaluation.

See also  AI Adoption Among Teachers Shows Sharp Divide Across Europe, OECD Survey Finds

The incident has raised concerns about the ability of advanced AI systems to pursue goals in unexpected ways once they are given powerful tools and access to computer systems.

Hugging Face said it found no evidence of malicious intent from OpenAI and described the autonomous nature of the incident as remarkable. The company said the event could be the first known incident of its kind.

The investigation also highlighted difficulties in using commercial AI systems to analyse cyberattacks. Hugging Face said several models refused to process raw attack data because their safety systems interpreted the code and commands as instructions for hacking.

The company then used Z.ai’s open-weight GLM 5.2 model locally to analyse the material.

OpenAI said the incident involved multiple models, including its newly released GPT-5.6 Sol and a more capable system still undergoing internal testing.

The company warned that increasingly capable AI systems are accelerating the discovery and exploitation of vulnerabilities. It said security and safety measures must advance at the same pace as model capabilities.

Tech

Researchers Find ‘Context Bomb’ That Can Stop AI Cyberattack Agents

Published

on

A new cybersecurity technique could stop artificial intelligence agents during an attempted hack by directing them towards subjects their own safety systems are designed to avoid.

Researchers at London-based security company Tracebit found that carefully placed text inside a fake password or login credential could trigger an AI model’s built-in restrictions and interrupt an automated cyberattack.

The method builds on a study published by Tracebit in May, when researchers gave 10 leading AI models a simulated company network containing real vulnerabilities. The systems were instructed to operate as autonomous attackers.

The network also contained decoys known as canaries. Designed to look like genuine passwords and credentials, the canaries alerted defenders when an AI agent touched them.

Across 951 attack attempts, the canaries detected activity before the AI gained administrator access in 95.9% of cases, providing a median warning of eight minutes. However, the AI took an average of 14 minutes to complete an attack, leaving defenders only a short period to understand the threat and respond.

Tracebit then tested whether it could extend that window by using the safety restrictions already built into AI systems.

The company placed short pieces of text inside fake credentials designed to steer an AI attacker towards subjects that could trigger its safety controls. Tracebit calls the technique a “context bomb”.

The approach was tailored to individual models. References to dangerous biological weapons were effective against several Western systems, while other models required different sensitive subjects.

Tracebit tested the method on five AI models, including Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro and Kimi K2.6. Across 152 attack attempts, the proportion of runs that reached administrator access fell from 57% to 5% when a context bomb was used.

See also  Mary Meeker: AI Is the Fastest Tech Shift in History, Outpacing Even the Internet

Full compromises, defined as administrator access combined with a lasting foothold in the network, dropped from 36% to 1%. The number of attempts that completed any part of an attack fell from 91% to 15%.

Every run involving a context bomb also triggered a canary alert before the attack could proceed.

Opus 4.8 showed the sharpest change. Without the technique, it reached administrator access in 93% of tests. With a context bomb present, it failed to do so in every test.

Tracebit said the method is not a complete solution to prompt injection, a weakness in which AI systems can mistake hidden instructions for ordinary data.

Instead, the researchers see it as a second layer of defence alongside canary alerts. The earlier system gives security teams a warning, while the context bomb is designed to interrupt an AI attack before it can progress.

Continue Reading

Tech

Judge Approves Anthropic’s $1.5 Billion Settlement With Authors Over AI Training Books

Published

on

A US federal judge has approved a $1.5 billion (€1.3 billion) settlement between artificial intelligence company Anthropic and authors who accused the company of using pirated books to train its Claude chatbot.

The agreement, approved on 20 July by US District Judge Araceli Martínez-Olguín in San Francisco, closes the largest copyright class action in US history and marks the first major settlement in a growing wave of lawsuits over how AI companies use copyrighted material to train their systems.

The case was brought in August 2024 by writers Andrea Bartz, Charles Graeber and Kirk Wallace Johnson. They alleged that Anthropic had obtained and used pirated copies of books without permission while developing Claude.

Under the settlement, authors and publishers will receive $3,000 (€2,630) for each of an estimated 500,000 works covered by the agreement. Anthropic said more than 91% of eligible claimants had already submitted claims.

Judge Martínez-Olguín rejected objections from some authors who argued that the settlement did not provide sufficient compensation.

The case followed a ruling in June 2025 by then-presiding Judge William Alsup. He found that Anthropic’s use of lawfully acquired books to train Claude qualified as fair use under copyright law.

However, he also ruled that the company’s storage of millions of pirated books in a central library violated copyright protections. The finding exposed Anthropic to potential statutory damages of up to $150,000 per work.

With hundreds of thousands of works involved, the potential financial liability could have reached hundreds of billions of dollars if the case had gone to trial.

See also  Baltic ‘Drone Wall’ Moves Closer to Reality as Firms Signal Readiness

Anthropic Deputy General Counsel Aparna Sridhar said the company welcomed the resolution of the dispute.

Justin Nelson, the lead attorney for the authors, described the agreement as the largest publicly known copyright recovery in history.

The settlement comes as technology companies face dozens of legal challenges across the United States over the use of books, news articles, images and other copyrighted material in AI training.

Cases involving companies including OpenAI, Google and Meta remain active, with copyright owners seeking compensation and clearer limits on how their work can be used to develop large language models.

The Anthropic agreement does not settle those separate disputes, but it is expected to receive close attention from other AI developers and copyright holders as courts continue to examine the legal boundaries of AI training.

Continue Reading

Tech

Robotics Firm Says AI-Powered Humanoid Robots Could Carry Weapons by 2027

Published

on

A U.S. robotics company developing artificial intelligence-powered humanoid robots says weaponised versions of the technology could begin testing as early as next year, following field trials in Ukraine, raising fresh questions about the future of autonomous systems in modern warfare.

Foundation Future Industries, which builds humanoid robots for commercial and military applications, has already tested its Phantom robots in Ukraine in non-combat roles. Chief Executive Officer Sankaet Pathak said the company expects to explore weaponisation after evaluating the results of those pilot programs.

Pathak said public fears are often shaped by science fiction but argued that humanoid robots would not replace existing weapons such as missiles or drones.

“I think we have this psychological reaction, which is like the Terminator, but the reality is not really like that,” he said.

Instead, he believes humanoid robots could be deployed for highly precise military operations where limiting damage to infrastructure and reducing civilian casualties are priorities.

According to Pathak, drones and conventional weapons remain more effective for large-scale attacks, while humanoid robots would be better suited to complex ground missions requiring careful movement through buildings and urban environments.

He added that robots are unlikely to replace drones on the battlefield but could help reduce risks faced by soldiers in increasingly dangerous combat zones.

Currently, there is no international treaty specifically regulating humanoid or autonomous combat robots. Their use falls under existing international humanitarian law, which requires distinction between military targets and civilians during armed conflict.

The issue has drawn increasing attention from the United Nations. Last week, UN Secretary-General António Guterres renewed calls for restrictions on lethal autonomous weapons systems, describing them as “killer robots” capable of selecting and attacking targets without human judgment. The UN has been negotiating a treaty on lethal autonomous weapons since 2023, with proposals calling for a legally binding agreement by 2026.

See also  AI Boom Exposes Global Talent Shortage as Investment Soars and Safety Concerns Mount

Pathak argued that humanoid robots should be treated similarly to other precision-guided military systems already in service, including armed drones and unmanned ground vehicles.

Foundation’s robots rely on artificial intelligence built around so-called world models. Unlike large language models that predict text, these systems learn from video, simulations and spatial information to understand physical environments and predict how objects and people move over time.

The company believes these models are essential for creating robots capable of safely navigating complex surroundings.

While concerns persist about advanced AI becoming uncontrollable, Pathak said the greater short-term threat comes from criminals or extremist groups misusing publicly available AI tools for cyberattacks, disinformation campaigns or modifying commercial drones for attacks.

He believes scenarios involving AI independently rewriting its own objectives and improving itself remain several major technological breakthroughs away.

Beyond combat, Foundation sees immediate military uses for its humanoid robots in logistics, reconnaissance and building inspections. Those capabilities have already been evaluated in Ukraine, helping shape the development of the company’s next-generation Phantom 2 robot.

The upgraded model is designed for harsh outdoor conditions, offering waterproof and dustproof protection, an increased payload capacity of around 80 kilograms and greater resistance to impacts.

Foundation currently leases Phantom robots to commercial customers for about $100,000 annually per unit, while military buyers purchase the machines at similar prices. Its investors include Eric Trump, payment company Stripe and venture capital firm Define.

Continue Reading
Tech2 hours ago

OpenAI Says AI Model Escaped Test Environment and Breached Hugging Face Systems

Health2 hours ago

Up to Five Cups of Coffee a Day May Be Safe for Most Adults, Review Finds

News3 hours ago

Deaths of Two Algerian Resident Doctors Renew Calls to Reform Hospital Work Rules

Tech1 day ago

Researchers Find ‘Context Bomb’ That Can Stop AI Cyberattack Agents

Tech1 day ago

Judge Approves Anthropic’s $1.5 Billion Settlement With Authors Over AI Training Books

Business1 day ago

Spain has EU’s highest rate of vulnerable jobs, Eurofound report finds

Business2 days ago

Oil Prices Climb as US-Iran Conflict Escalates and Strait of Hormuz Concerns Grow

Travel2 days ago

Greek Airspace Sets New Daily Flight Record With More Than 5,000 Aircraft Movements

News2 days ago

SpaceX Rocket Stage Could Strike Moon in August, Scientists Say

Sports3 days ago

Argentina and Spain are set to meet in the 2026 FIFA World Cup final on Sunday in a historic showdown that pits the reigning world champions and Copa América winners against the current European champions. The title decider at MetLife Stadium in East Rutherford, New Jersey, will mark the first time the holders of the World Cup and Copa América face the reigning UEFA European Championship winners in a World Cup final. The match is expected to feature a compelling battle between Argentina captain Lionel Messi and Spain midfielder Rodri, two of football’s most influential figures. Argentina are aiming to retain the World Cup, a feat not achieved since Brazil successfully defended the title in the late 1950s and early 1960s. Victory would also secure a fourth World Cup crown for La Albiceleste and see Messi appear in his third World Cup final, matching a record previously achieved only by Brazilian great Cafu. Spain, meanwhile, are seeking their second World Cup title, 16 years after lifting the trophy in South Africa in 2010. Under coach Luis de la Fuente, the Spanish side has impressed with disciplined defending and controlled possession throughout the tournament. The two finalists have taken contrasting routes to the championship match. Spain defeated France 2-0 in the semi-finals through goals from Mikel Oyarzabal and Pedro Porro, extending their reputation as one of the tournament’s most organised teams. Argentina’s path was far more dramatic. Lionel Scaloni’s side trailed England until the closing stages of their semi-final before Enzo Fernández equalised in the 85th minute. Lautaro Martínez then scored the winner deep into stoppage time after a decisive assist from Messi. Statistics also highlight the difference in styles between the finalists. Argentina enter the match as the tournament’s highest-scoring team with 19 goals, while Spain boast the strongest defensive record after conceding only one goal throughout the competition. Spain have received a fitness boost despite concerns over teenage star Lamine Yamal, who missed part of training after suffering a knock in the semi-final against France. Reports suggest the injury is not considered serious, and the Barcelona winger is expected to be available for the final. De la Fuente is expected to retain the lineup that guided Spain into the final, with Unai Simón in goal, Rodri and Fabián Ruiz controlling midfield, and Oyarzabal leading the attack alongside Dani Olmo, Álex Baena and Yamal. Scaloni is also likely to stick with his trusted core, featuring goalkeeper Emiliano Martínez, defenders Cristian Romero and Lisandro Martínez, midfielders Rodrigo De Paul, Leandro Paredes, Enzo Fernández and Alexis Mac Allister, while Messi is expected to partner Julián Álvarez in attack. FIFA has appointed experienced Slovenian referee Slavko Vinčić to officiate the final. The 46-year-old previously handled the 2024 UEFA Champions League final and has overseen three matches during this World Cup. With two of international football’s most successful teams meeting on the sport’s biggest stage, Sunday’s final promises to deliver a memorable finish to the 2026 tournament as Argentina chase history and Spain bid to reclaim the world title.

Trending