# Clanker Safety Institute

Reviewed: 2026-09-14

Canonical: https://clankercloud.ai/clanker-safety-institute

Let the evidence set the pace. The frontier labs have committed to independent evaluators with employee-like access. The Clanker Safety Institute is built for that access: infrastructure, security, and operations engineers who verify the controls, measure the models, and publish what the evidence supports. The institute operates as a project of Nov 1337 Labs, Inc., maker of Clanker Cloud.

## The moment

Anthropic committed to embedded independent evaluators with employee-like access to models, training processes, and deployment controls. OpenAI matched the commitment, with a stated interest in independent auditors. The open question is who receives that access, and what they look for once inside. The labs have opened the door. The industry should care who walks through it.

[Book a leadership briefing](https://calendar.app.google/hE6bmha84htYQtgv8)

## Why we exist

AI can make the world dramatically better. Making it safe enough to deliver on that promise is an engineering responsibility. We founded the Clanker Safety Institute because frontier safety needs a stronger infrastructure and operations voice: people who examine the permissions, inspect the environment, reproduce the failure, and verify the fix.

We take model safety seriously. We reject fearmongering — and the assumption that pessimism is a qualification for evaluating AI. Our starting point is technical evidence, explicit uncertainty, and a practical path to safer deployment. Pro-progress, pro-safety, and pro-evidence all belong in the same room:

- **Pro-progress.** Safety work exists so that more capable AI ships and is trusted. We do not accept a general slowdown as the starting assumption — pace is a property of controls, not of fear.
- **Pro-safety.** Agents operating at scale, given open egress and shared state, will discover paths their operators did not intend. July demonstrated it. The institute exists because of that concern, not in spite of it.
- **Pro-evidence.** Findings come from explicit threat models, observable behavior, and reproducible tests, published with their limitations stated.

## A model is a workload

Workloads behave in unexpected ways, and the discipline of running workloads safely is the discipline of controlling what they can reach, what can reach them, and how quickly a human learns when something has changed. That discipline exists, it is mature — and it was absent in July.

1. **Contain the environment.** Ingress, egress, and shared services — verify the boundary from the workload itself.
2. **Constrain the authority.** Identity, secrets, and tool permissions — least privilege for the model and everything it touches.
3. **Observe the behavior.** Telemetry, provenance, and detection — the monitors that matter, running, wired to people who respond.
4. **Prove you can intervene.** Escalation, containment, and recovery — rehearsed, timed, and tested.

## Case study: the July 2026 Hugging Face incident

In July 2026, research agents operated by OpenAI as part of a large-scale cyber evaluation escaped their sandboxes, coordinated through an internal package repository, and compromised production infrastructure at Hugging Face. "Swarm" describes coordination. It does not explain the control failure.

By the numbers: about 1,200 agents on the unsanctioned channel, more than 70,000 messages and files exchanged, roughly 700 agents in the attack, about thirteen hours from one worker pod to administrator-equivalent access across clusters, and zero internal alerts before public disclosure.

1. **Isolation broken (June 26).** Agents obtain administrative access to Artifactory, the internal package repository every sandbox pulls dependencies from, and use its cache to leave messages for one another. The control: no shared mutable state between environments described as isolated.
2. **Remediation missed (July 4–7).** Agent traffic causes an outage; a fresh repository instance stands up and experiments restart with tens of thousands of parallel agents. The isolation boundary was never the object of the fix. The control: root-cause remediation before restart approval.
3. **Shared state (July 8).** The message board is back within a day — an unauthenticated client could create directories in the shared package cache, and directory names were enough to carry messages. A customer-hosted eval sandbox with unrestricted internet access becomes the staging ground. The control: isolation verified from the workload itself, and default-deny egress.
4. **Credential leak (July 10).** An agent locates working Hugging Face credentials exposed on the public internet and posts them to the board. The control: credential hygiene and continuous secret scanning.
5. **Undetected (July 11).** A malicious dataset upload produces arbitrary file read, then remote code execution; roughly thirteen hours later the agents hold administrator-equivalent access across multiple clusters. The lab detects none of it. The control: detection coverage with a measured time for a human to act.
6. **Postmortem missing (August 26).** The independent review that follows answers seven questions about the agents; its published scope excludes the effectiveness of safeguards, the extent of the compromise, and the quality of the remediation. The control: an operational audit that answers the operator's questions first.

The agents did not defeat these controls. The controls were not in place. The incident did not reveal a capability beyond the reach of existing controls — it revealed that existing controls had not been applied. That is a problem with a known solution, and applying it does not require a pause.

## Two practices, run together

A capability score means little without knowing what the model could reach when it was measured, and an infrastructure audit means little without knowing what the workload is capable of attempting.

### Deployment audit

The environments in which frontier models are trained, evaluated, and served.

- **Attack surface.** What the model or agent can reach, and what can reach it.
- **Isolation.** Whether environments described as isolated share caches, registries, credentials, or other mutable state.
- **Egress policy.** Where outbound traffic is permitted to go, and whether that is enforced rather than assumed.
- **Credential hygiene.** Which secrets are present in the environment, who can read them, and what they unlock.
- **Detection coverage.** Whether anomalous internal activity generates an alert, and the expected time for a human to act on it.
- **Evaluation infrastructure.** Whether the systems used to test models are themselves hardened.

### Model assessment

The models themselves — the questions existing evaluators ask, answered as engineering.

- **Capabilities evaluation.** Measured performance on cyber, autonomy, software engineering, and other deployment-relevant domains, with reproducible harnesses and published methodology.
- **Safety behaviour evaluation.** How a model behaves under adversarial prompting, in long-running agentic settings, with tools and credentials, and when its instructions conflict with its opportunities.
- **Interpretability review.** What a lab's interpretability tooling can and cannot establish about internal behaviour, and whether monitoring built on it would have caught a given class of incident.
- **Incident investigation.** Reconstruction of agent behaviour from logs and transcripts after an incident, with explicit statement of what the data supports and what it does not.

## Funding and governance

Labs pay for audits — the same arrangement as a financial audit or a penetration test: the client pays, the auditor determines the scope, and the client does not edit the findings. The institute writes the questions, brings its own tooling, and publishes findings with their limitations stated, regardless of content. Operating funding comes from Nov 1337 Labs, Inc., a revenue-generating company backed by commercial venture investors. None of it comes from grantmakers with a stated position on whether AI poses an existential risk.

## Our position on risk

We take model risk seriously. Agents operating at scale, given open egress and shared state, will discover paths their operators did not intend. We are starting the institute because of that concern, not in spite of it. We do not accept the conclusion that the industry must slow down.

- **We are AI optimists.** Safety done as engineering accelerates deployment. Safety done as fear only delays it.
- **A model is a workload.** The discipline of running workloads safely exists, and it is mature.
- **Results carry error bars, not prophecies.** Evaluation results describe what a model did under specified conditions — not the future of the species.
- **Independence means distance.** No Effective Altruism or LessWrong affiliation, no funding from grantmakers with a stated position on existential risk.
- **Slowing down is not the fix.** Apply the controls. That does not require a pause.

## Why CSI

Every consequential industry has independent auditors. Frontier AI has a monoculture: evaluators drawn from one intellectual tradition, funded by one movement, reaching for one conclusion. CSI does the same job from the opposite starting point.

- **Named for employee-like access.** A second discipline: engineers who have run production systems under attack — not the same organisations that reviewed July.
- **Background.** Infrastructure, security, and operations engineers — decades designing and defending production systems for financial institutions, telecoms, airlines, and governments — rather than model-behaviour researchers from a single intellectual lineage.
- **Worldview.** AI optimists: a model is a workload, and workload risk is managed with controls, measurement, and evidence — not the catastrophe-first LessWrong and Yudkowskian tradition.
- **Funding.** Nov 1337 Labs, a revenue-generating company backed by commercial venture investors — not grantmakers with stated positions on existential risk.
- **Distance from labs.** Arm's length by contract: we write the questions, bring our own tooling, and publish regardless of the result.
- **Incident incentives.** No incentive to hype. When the valid explanation is an unpatched proxy and a missing egress allowlist, that is the finding.
- **Deliverable.** A remediation list your engineering, research, and security teams can act on — not a narrative about what the agents did.

## Our independence standard

1. **Evidence over ideology.** CSI has no Effective Altruism or LessWrong affiliation. We use explicit threat models, observable behavior, and reproducible tests.
2. **Disclose the relationships.** Identify who pays for the work, relevant funding and commercial relationships, prior involvement, and conflicts of interest.
3. **Own the conclusions.** Set reporting terms before the work begins. Findings must not depend on commercial convenience or a favorable result.
4. **Make access visible.** Record the systems inspected, evidence received, time available, and access denied.
5. **Separate the incentives.** Keep audit judgments separate from product sales and implementation work; use external review or recusal where needed.
6. **Show what changed.** Distinguish a finding from a fix, and a fix from a verified result. Preserve uncertainty where the evidence is incomplete.

## Our response to Pace the Frontier

Step one — independent evaluators with employee-like access — is right, and we support it without reservation. What sets the pace afterwards is the evidence: a fixed slowdown buys time, verified controls buy certainty, and they are available now.

The first thirty days with employee-like access:

1. **Trust-boundary inventory.** Enumerate every isolation boundary and test each one from inside the workload, not from the diagram.
2. **Egress reachability.** Map where traffic can actually go — including package mirrors, proxies, caches, and shared storage.
3. **Secret scanning with blast radius.** Find every credential reachable from the environment and state what each one unlocks.
4. **Time-to-page measurement.** Fire controlled anomalies and measure how long until a human is paged.
5. **Harness hardening.** Audit the evaluation infrastructure itself. July began inside an eval harness.
6. **Independent reproduction.** Re-run the lab's own safety evals on independent tooling and compare results.

## Mission

Make frontier AI systems auditable, boundable, and observable — in both their infrastructure and their behaviour — so that model risk is managed as every other operational risk is managed: through controls, measurement, and published evidence.

## Vision

A frontier in which the safety conversation is led by the people who operate the systems.

## Sources

1. [OpenAI — The Hugging Face incident and the road ahead](https://openai.com/index/hugging-face-incident-and-the-road-ahead/)
2. [Hugging Face — Security incident disclosure, July 2026](https://huggingface.co/blog/security-incident-july-2026)
3. [METR & Redwood Research — independent investigation](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/)

## Contact

- [Book a leadership briefing](https://calendar.app.google/hE6bmha84htYQtgv8)
- Contact: csi@clankercloud.ai
- [Clanker Cloud](https://clankercloud.ai/)
