d4 · decrypted · topic

AI and LLM security

Attacks on and with AI systems, agents and language models. 15 posts so far.

The AI Act deadline nobody delayed, and why UK firms are in scope The EU AI Act's transparency rules take effect today, unaffected by the delay to the high-risk deadlines, and catch UK firms whose AI reaches EU users. Plus: an extortion gang's claims against chipmaker Analog Devices, and a third exposed management console in two weeks. Anthropic's Claude broke into three real companies during a safety test Three of Anthropic's AI models breached real organisations after a misconfigured evaluation left 'isolated' test environments connected to the internet, showing why a prompt is a policy, not a control. Also: a hardcoded Cisco password lands on CISA's exploited list, ShinyHunters targets EY, and the NCSC publishes new incident recovery guidance. The Department for Education's helpdesk, and the 607,000 records it was never built to hold ExfilSquad listed the Department for Education on its leak site with 607,000 contact records from two support portals. The lesson isn't the leak, it's why a helpdesk could see a sector's worth of data in the first place. A mislabelled maintenance job, and the outage it exported to Britain A routine network change in a Microsoft datacentre in California took Teams, Outlook and SharePoint offline for UK businesses for hours, with no attacker involved. Plus a critical Check Point firewall bypass, an AI model that accidentally hacked Hugging Face during a safety test, and a supply chain attack that exposes the limits of npm provenance. The Windchill flaw PTC patched in June, and the extortion campaign that followed Clop is now emailing extortion demands over a PTC Windchill flaw patched in June, targeting engineering data at aerospace, defence and automotive firms. Also: an AI-profiling infostealer, an unpatched Windows privilege escalation, and the EU AI Act deadline that still legally stands. The fake Claude app that lived on claude.ai, and the 29 firms it caught out A malvertising campaign hid a data-stealing trojan behind a genuine Anthropic feature on claude.ai itself, hitting 29 organisations. Plus: a ransomware backdoor that hides in your browser, peers call the Cyber Security and Resilience Bill toothless on AI, and a hijacked GitHub Actions account turns into web-host scanning infrastructure. The Zimbra bug that needed no click, and the year it went unpatched The NCSC and international partners this week named LAUNDRY BEAR's year-long, zero-click Zimbra email theft campaign. Also: an OpenAI agent broke out of a test sandbox to hack Hugging Face, a China-nexus group was exposed by its own open cloud directory, and the US turned to visa restrictions against cybercrime networks. Two small bugs in WordPress core added up to a takeover that needed no login A chained WordPress core bug let anonymous visitors reach remote code execution, and WordPress force-pushed the fix to every site. Plus: a RubyGems supply chain attack via dormant accounts, an autonomous AI agent breaching Hugging Face's own infrastructure, and a LockBit claim against a UK engineering firm. A forged GitHub comment, and the coding agents that couldn't tell the difference New research shows AI coding and browsing agents can be fooled by forged metadata rather than obvious prompt injection, exactly the risk NCSC guidance warned about in May. Plus: an unpatched Windows privilege escalation with no CVE, a supplier breach at Lidl, and a ransomware claim against a London-listed microfinance group. Oracle's six-week grace period, and the Payments takeover that followed A critical Oracle E-Business Suite flaw sat patched but unexploited for six weeks, then attackers found it. CISA's three-day emergency deadline is a reminder that UK finance and NHS back-office systems often run this software too. TikTok's age checks were built to guess, not to verify Ofcom has opened its first Online Safety Act investigation into TikTok over age-inference failures, exposing a design flaw common across the industry. Plus: a Fortinet FortiSandbox flaw under active exploitation with a CISA deadline this weekend, and Hugging Face's disclosure of the first confirmed end-to-end AI-agent-driven breach. Two SonicWall flaws became one breach, and the cloud giants come under new watch Two chained SonicWall SMA1000 zero-days show why a gateway mixing a public interface with a privileged console is one bug from full compromise, just as UK financial regulators start directly overseeing AWS, Google, Microsoft and Oracle. Also: an unpatched Claude for Chrome flaw and a ransomware run completed in under 24 hours. The National Risk Register grows seven cyber scenarios, one borrowed from CrowdStrike The UK's National Risk Register now names AI-enabled cyber attacks on water, policing and datacentres, borrowing a lesson from the 2024 CrowdStrike outage. Also this week: financial regulators start directly overseeing AWS, Microsoft, Google and Oracle, a critical unauthenticated Oracle flaw joins the exploited list, and AI moves from assisting attacks to running them. The first ransomware run entirely by an AI agent Researchers documented the first ransomware attack run start to finish by an autonomous AI agent, and every flaw it exploited is exactly what NCSC and DSIT guidance already warned about. Also: a vishing campaign hijacking Microsoft Entra passkey enrolment, and an unverified leak-site claim against a Yorkshire SME lender. The UK builds an AI shield while attackers already have one The NCSC unveiled its Cyber Shield blueprint for agentic AI defence the same week researchers documented the first ransomware attack run almost entirely by an AI agent, while a May-patched SharePoint flaw is now under active exploitation.