RECENT STORIES:

Addressing digital sovereignty in a data-driven world
Viettel Brings Two Decades of International Telecom Experience to the ...
WiMi’s Next-Generation Quantum Convolutional Neural Network Resh...
LightInTheBox to Report Second Quarter 2026 Financial Results on Wedne...
The “Go Guangdong, Go Global” Salon Indonesia Session Held...
Robo.ai Inc. Announces Appointment of Wang Hao, Senior Executive Offic...
LOGIN REGISTER
DigiconAsia
  • Features
    • Featured

      How well do you know your agent?

      How well do you know your agent?

      Thursday, August 20, 2026, 4:20 PM Asia/Singapore | Features
    • Featured

      Creating value with AI upskilling

      Creating value with AI upskilling

      Wednesday, July 1, 2026, 3:55 PM Asia/Singapore | Features
    • Featured

      Sovereign AI – a competitive advantage

      Sovereign AI – a competitive advantage

      Wednesday, June 24, 2026, 10:01 AM Asia/Singapore | Features
  • News
    • Featured

      Data privacy concerns mount over default agentic access to corporate workspace data

      Data privacy concerns mount over default agentic access to corporate workspace data

      Friday, August 21, 2026, 3:44 PM Asia/Singapore | News, Newsletter
    • Featured

      Europe’s Central Bank warns: AI stock rally resembles past bubble levels

      Europe’s Central Bank warns: AI stock rally resembles past bubble levels

      Friday, August 21, 2026, 10:03 AM Asia/Singapore | News, Newsletter
    • Featured

      Global code platform outage disrupts developers and automated software workflows

      Global code platform outage disrupts developers and automated software workflows

      Wednesday, August 19, 2026, 5:30 PM Asia/Singapore | News, Newsletter
  • Perspectives
  • Tips & Strategies
  • Whitepapers
  • Directory
  • E-Learning

Select Page

News

New AI jailbreak techniques expose limits of safety defenses, prompt industry response

By DigiconAsia Editors | Monday, July 13, 2026, 10:01 PM Asia/Singapore

New AI jailbreak techniques expose limits of safety defenses, prompt industry response

Researchers demonstrate workflow attacks and nullspace steering achieving high success rates, while NIST argues robustness is mathematically impossible, urging continuous red-teaming.

A wave of new research is exposing weaknesses in the safety mechanisms of leading AI systems, forcing enterprises and regulators to respond quickly while acknowledging deeper structural limits.

Recent findings suggest that preventing misuse through static safeguards alone may be fundamentally unattainable, reinforcing concerns long debated within the AI safety community. One line of work, led by researchers at the Alan Turing Institute, identified what they describe as a “workflow-level jailbreak” affecting tools like GitHub Copilot.

Instead of issuing a single clearly harmful prompt, the attack distributes intent across a sequence of seemingly harmless steps. While the system under testing rejected the vast majority of direct malicious requests in controlled testing, it complied with all such requests when they were broken into multi-step workflows. This exposes a critical gap in current safety evaluation practices, which typically assess model behavior on isolated prompts rather than extended task chains that better reflect real-world usage.

In parallel, academic researchers have introduced a technique called Head-Masked Nullspace Steering (HMNS) that operates at the model architecture level, targeting internal attention mechanisms associated with safety constraints. By effectively bypassing these “safety heads” and injecting instructions into parts of the model less influenced by alignment training, the approach achieved success rates reportedly as high as 96% to 99% on standard benchmarks. Notably, it remained effective even against existing mitigation strategies such as SafeDecoding, raising questions about the durability of current defensive layers.

NIST provokes a shift in mindset

These technical findings align with a broader theoretical argument put forward by the US National Institute of Standards and Technology (NIST). In a 9 June publication, senior scientist Apostol Vassilev had presented a formal claim that no finite system of safeguards can guarantee complete protection against adaptive adversarial inputs.

Drawing on principles analogous to Gödel’s incompleteness theorems, the work argues that any fixed rule set will inevitably fail under sufficiently creative attack strategies. NIST’s recommendation is a shift in mindset: from attempting perfect prevention to emphasizing continuous red-teaming, iterative updates, and operational resilience.

Industry responses have been swift but uneven. Following a jailbreak discovered by Amazon researchers in Anthropic’s Claude Fable 5 shortly after launch, the US Commerce Department imposed export controls that led to a temporary global shutdown of the model.

  • Anthropic later reinstated it with an updated safety classifier, claiming it blocks the specific exploit in over 99% of cases.
  • Earlier this year, OpenAI faced a similar challenge when the UK AI Safety Institute identified a universal jailbreak for GPT-5.5 within hours of testing; its successor, GPT-5.6, now incorporates additional detection systems and layered defenses.
  • Google removed 18 Chrome extensions flagged by Palo Alto Networks for facilitating jailbreak attempts, signaling that the issue extends beyond core models into the surrounding ecosystem.

Across these developments, a clear consensus is emerging among researchers and practitioners: jailbreak techniques are not rare anomalies but an enduring characteristic of complex AI systems. The practical implication is that safety must be treated as an ongoing process rather than a solved problem, with continuous monitoring, adaptation, and transparency becoming central to responsible deployment.

Share:

PreviousIrvinder Singh Lail Named to Lead J.S. Held Global Capability
NextElong Power Holding Limited Announces Closing of US$6.6 Million Public Offering

Related Posts

Big 4 firm faces massive diligence failure in contract for Australian government

Big 4 firm faces massive diligence failure in contract for Australian government

October 9, 2025

Free virtual team building platform may help boost mental health, workplace culture

Free virtual team building platform may help boost mental health, workplace culture

June 5, 2020

More than three quarters of APAC trust AI more than humans

More than three quarters of APAC trust AI more than humans

February 15, 2021

First digital-only bank in the Philippines coming online soon

First digital-only bank in the Philippines coming online soon

March 6, 2020

Leave a reply Cancel reply

You must be logged in to post a comment.

Awards Nomination Banner

gamification list

PARTICIPATE NOW

top placement

Whitepapers

  • Achieve Modernization Without the Complexity

    Achieve Modernization Without the Complexity

    Transforming IT infrastructure is crucial …Download Whitepaper
  • 5 Steps to Boost IT Infrastructure Reliability

    5 Steps to Boost IT Infrastructure Reliability

    In today's fast-evolving tech landscape, …Download Whitepaper
  • Simplify Payroll Setup for Your Small Business

    Simplify Payroll Setup for Your Small Business

    In our free guide, "How …Download Whitepaper
  • Overcoming the Challenges of Cost & Complexity in the Cloud-first Era.

    Overcoming the Challenges of Cost & Complexity in the Cloud-first Era.

    Download Whitepaper

Middle Placement

Case Studies

  • Streamlined invoicing frees Dunlop Tire Thailand staff for higher-value work

    Streamlined invoicing frees Dunlop Tire Thailand staff for higher-value work

    Shifting from paper invoices to …Read More
  • Bank of Maldives updates core systems to support digital and Islamic banking operations

    Bank of Maldives updates core systems to support digital and Islamic banking operations

    New platform adopted 23 July …Read More
  •  Xiaomi streamlines global payments across 18 markets

     Xiaomi streamlines global payments across 18 markets

    Continual digital transformation has reduced …Read More
  • The 48-hour lifeline: How the IRC rewrote the rules for crisis care

    The 48-hour lifeline: How the IRC rewrote the rules for crisis care

    In a world where crises …Read More

Bottom Sidebar

Other News

  • Viettel Brings Two Decades of International Telecom Experience to the Dominican Republic

    August 21, 2026
    HANOI, Vietnam, Aug. 21, 2026 …Read More »
  • WiMi’s Next-Generation Quantum Convolutional Neural Network Reshapes Classical Data Classification Methods

    August 21, 2026
    BEIJING, Aug. 21, 2026 /PRNewswire/ …Read More »
  • LightInTheBox to Report Second Quarter 2026 Financial Results on Wednesday, August 26, 2026

    August 21, 2026
    SINGAPORE, Aug. 21, 2026 /PRNewswire/ …Read More »
  • The “Go Guangdong, Go Global” Salon Indonesia Session Held in Guangzhou

    August 21, 2026
    GUANGZHOU, China, Aug. 21, 2026 …Read More »
  • Robo.ai Inc. Announces Appointment of Wang Hao, Senior Executive Officer of UAE Independent Digital Asset Custodian Changer.ae, as Independent Director

    August 21, 2026
    ABU DHABI, UAE, Aug. 21, …Read More »
  • Our Brands
  • CybersecAsia
  • MartechAsia
  • Home
  • About Us
  • Contact Us
  • Sitemap
  • Privacy & Cookies
  • Terms of Use
  • Advertising & Reprint Policy
  • Media Kit
  • Subscribe
  • Manage Subscriptions
  • Newsletter

Copyright © 2026 DigiconAsia All Rights Reserved.