decrypted · 4 october 2026 · vulnerabilities and patching · supply chain · ai and llm security
Half a million working keys sit in public code, and some have been there since 2009
A security firm has found more than half a million working passwords, keys and tokens sitting in public code repositories, and the most uncomfortable number is not the total but the age. Truffle Security scanned 224 million public GitHub repositories and found 543,699 unique credentials that still authenticated in July 2026. The median one had been exposed for 784 days. The oldest dated from 2009 and still worked.
What the researchers did
The corpus was The Stack v3, a dataset assembled for training large language models. Truffle matched known credential patterns, then tested each hit against the issuing provider on 27 and 28 July 2026. A credential only counted if the provider confirmed it was live. Those 543,699 secrets appeared 1,103,438 times, because the same secret is often copied around.
Think of a hotel that never recodes its room locks. Guests leave, keys go missing, and nobody knows which doors they still open. Committing a secret to a public repository is leaving a key in a car park. Deleting the file afterwards changes little, because the history keeps a copy and the lock is still the same lock.
Why some keys died and others did not
The survival rates are the real lesson. Truffle reported that almost no npm tokens (0.001%) and few GitHub tokens (0.36%) still worked, because those providers run revocation programmes. By contrast, 54% of Google Cloud service account credentials and 88% of Postgres connection strings were still live. Nobody was watching for those, and nothing expired them.
GitHub made push protection the default from 29 February 2024. It blocks known secret formats at the moment of commit, and Truffle estimates it cut leakage in covered categories by about 53%. But 51.8% of the live credentials were in a shape that a default public repository will still accept, and the feature cannot reach commits made before it existed.
The Secure by Design lesson
For years the burden has sat on the individual developer, who is asked never to make a mistake. A secure by design platform assumes the mistake will happen and limits what it costs. Three decisions would have helped:
- Issue short-lived credentials by default, so a leaked key is dead within hours rather than after 784 days.
- Prefer workload identity, where a system proves who it is without holding a long-lived secret at all.
- Have the issuer detect and revoke leaked secrets automatically, as npm and GitHub do for their tokens.
The cost is modest: some plumbing, and losing one key that works forever.
What UK organisations should do
This is opinion, not advice. Treat anything ever committed to a public repository as compromised, rotate it, and scan the history rather than only new commits, because Truffle is explicit that earlier commits fall outside push protection. Ask the agencies and suppliers who write code for you whether their secrets expire. Where a cloud provider offers short-lived tokens, write them into procurement requirements rather than treating them as an optional extra.
The training-data angle matters too. A key copied into a dataset built for AI models is a copy nobody can recall. If personal data sits behind a leaked database connection string, that is a data protection problem as well as a security one.
Also this week
MetaMask investigates an infrastructure incident. MetaMask said on 1 October that part of its infrastructure had been affected, and that there was no indication wallets or customer funds were hit. It is exiting affected Ethereum validators in its non-custodial staking operation as a precaution. Lido Finance said the last exits were expected by the end of 7 October, with likely foregone rewards and possible downtime penalties. MetaMask has not said what was compromised or whether data was accessed, so "no impact" remains the company's claim, not a finding. MetaMask says it does not hold clients' withdrawal keys, which is the design choice that limits the damage.
OpenAI's agents face scrutiny over scraping. The Record reports that Asymmetric Security, a forensics firm, says OpenAI software scraped data from dozens of websites, in some cases reaching staging environments and creating accounts. OpenAI told the outlet it was routine research using public information, and that it is investigating. Those are competing accounts and neither is independently established here. In the related Australian Medicare case, researchers cited by The Record found the site's archived code sent visitors to an unauthenticated guest endpoint. An agent that treats locks as puzzles is a risk, but so is a door left on the latch. UK organisations deploying agents should decide what they may touch before switching them on.
Sources
- Truffle Security: 543,699 exposed GitHub credentials nobody revoked
- BleepingComputer: Over 543,000 valid credentials exposed in public GitHub repositories
- MetaMask: security incident update
- BleepingComputer: Metamask discloses security incident affecting its infrastructure
- The Record: OpenAI software attempted to secretly scrape data from dozens of websites
- The Record: OpenAI agent and the Australian Medicare portal
Want a second pair of eyes on how your team handles secrets and supplier code? get in touch.
More like this
- An MCP flaw lets a rogue tool server steal an AI agent's login keys 29 september 2026
- An AI agent wiped Azure storage in seven minutes, using a password left on GitHub 28 september 2026
- A wormable DNS flaw headlines Microsoft's biggest Patch Tuesday yet 10 september 2026
Get the next post by email: subscribe to Decrypted. Double opt-in, unsubscribe any time, or take the RSS feed.
Prefer to listen? Decrypted on Apple Podcasts, or paste the podcast feed into any app.