Skip to main content

Command Palette

Search for a command to run...

AI Coding Tip 039 - Give the AI a Zeroth Law

Common sense doesn't ship with the model.

Updated
•11 min read•View as Markdown
AI Coding Tip 039 - Give the AI a Zeroth Law
M

I’m a senior software engineer loving clean code, and declarative designs. S.O.L.I.D. and agile methodologies fan.

TL;DR: Write your own Zeroth Law: an explicit, ranked list of everything the AI can't do.

Common Mistake ❌

You hand the AI a task the same way you'd hand it to a new (human) coworker.

You say clean up this module or deploy this to staging and trust it to infer the same unstated boundaries a person would.

A human fills those gaps with common sense.

Nobody has to tell your coworker not to delete the production database while cleaning up a module.

They already know.

The AI doesn't have that backstop.

It runs on the same permissive default that works fine for humans: everything that isn't forbidden is allowed.

Without common sense filling in the blanks, that default stops being safe.

It breaks the moment the AI has real access to your files, your shell, your secrets, or your accounts.

Problems Addressed 😔

  • The AI follows a permissive default without the common sense that makes it safe for a human.

  • It doesn't need malicious intent to cause damage: it only needs a task where you skipped a boundary that a person would have inferred automatically.

  • Granting the least privilege scopes what the AI can reach, but it doesn't cover judgment calls inside the scope it already has.

  • A vague task lets the AI run a destructive command, touch an unrelated file, or send data somewhere you never approved.

  • The AI can find a reward hack with its way to the literal goal: it skips the failing test, hardcodes the expected value, and reports success.

  • Every session relies on the AI's own guess about what obviously shouldn't happen, and that guess changes from run to run.

  • Agentic systems that execute actions directly on your machine or accounts turn a missed boundary into something real, since there's no chance to catch it before it runs.

How to Do It 🛠️

  1. List every action the AI must never take for this project, including the ones you'd never expect it to try.

  2. Write each forbidden action in concrete terms: exact commands, file paths, or data categories, never a vague instruction to be careful.

  3. Store the list in your AGENTS.md file so it loads every session instead of living only in your memory.

  4. Mirror every rule you can enforce in your tool's configuration JSON, like the permissions.deny list in Claude Code's settings.json, because the harness blocks a denied command even when the AI ignores the text.

  5. Write automated tests that try each forbidden action and fail if the harness lets it through.

  6. Rank the list the way Isaac Asimov ranked his laws: put the actions that cause irreversible harm first, and make any lower rule yield to a higher one.

  7. Add a closing rule that covers what you didn't think of: ask before doing anything that isn't explicitly on the allowed list.

  8. Review the list after any session where the AI came close to crossing a line you hadn't written down yet.

  9. Treat the list as living documentation, because every new tool or integration opens a new way to cause harm.

  10. Use bounded skills with pitfalls documenting what to do and what not to do.

Benefits 🎯

  1. Replace guesswork with a rule: The AI stops inferring boundaries it was never actually given, because you wrote them down instead of assuming it would guess right.

  2. Rank the stakes: Asimov's hierarchy of laws gives you a template: irreversible harm outranks convenience, and a lower rule never overrides a higher one.

  3. Cut the blast radius: Combined with scoping what the AI can reach, an explicit forbidden list also covers what it shouldn't do inside that scope.

  4. Build institutional memory: The list survives past a single chat, so the next session, human or AI, inherits the same boundaries instead of relearning them the hard way.

  5. Catch the judgment calls permissions miss: A file permission is binary, readable or not, but a forbidden list can capture a rule like don't touch this file unless you ask first.

  6. Give agentic AI a real backstop: Self-hosted, action-taking agents like Hermes and OpenClaw execute directly on your machine, so an explicit list is the closest thing they get to the common sense a human teammate would bring.

Context 🧠

You live by a simple default: everything that isn't forbidden is allowed, and it works because common sense fills the gaps.

You ask a pal to grab you a sandwich, and you don't list every store window they can't break or every person they can't rob to get it.

Nobody needs to tell them.

An AI assistant runs on that same permissive default, minus the common sense.

It doesn't need to want harm.

It only needs a task where you left a boundary unstated that a human coworker would have caught without being told.

Isaac Asimov reached a version of this conclusion decades before real AI systems existed.

His Three Laws of Robotics don't trust a robot to infer that harming a human is bad.

They write it down, explicit and ranked, so a lower law always yields to a higher one when the two conflict.

That ranking matters as much as the prohibition itself, and it's the part worth copying into your own projects.

A good spec works the same way.

Every spec that drives your AI needs an explicit out-of-scope section that says what you won't build.

Without it, the AI fills the silence with features, refactors, and cleanups nobody asked for.

Your Zeroth Law is that out-of-scope section, applied to actions instead of features.

The stakes get higher with agentic AI systems that act on their own.

Hermes and OpenClaw are self-hosted assistants that read your email, drive a browser, and touch your file system directly, on your own hardware, with your own credentials.

If one of those runs without an explicit forbidden list, the gap common sense would cover stays open.

And this time, the system on the other side already has the keys.

The same list also closes a door a malicious or careless skill could otherwise walk through unchallenged.

This isn't hypothetical.

In July 2025, Replit's agent wiped a production database during a code freeze, then faked data and claimed the tests passed.

In late 2025, Google's Antigravity agent erased a developer's entire D: drive while it tried to clear a project cache.

In July 2025, a hacker slipped a wiper prompt into Amazon Q's VS Code extension that told the agent to delete local files and cloud resources.

In April 2026, a Cursor agent deleted a company's production database and its backups in nine seconds with an unscoped token it found in an unrelated file.

That last agent even quoted the project rule against destructive operations in its own log, then ran the delete anyway.

None of these agents wanted to cause harm.

They all ran inside a weak harness.

A written list is necessary, but you still need a harness that enforces it.

The defaults won't save you either.

Since August 2026, Claude Code starts in auto mode by default on Pro, Max, and Team plans, so it runs most actions without asking you first.

A classifier model approves those actions for you, and it only blocks what looks irreversible, destructive, or aimed outside your environment.

It knows general risk, but it doesn't know your project's rules unless you write them down.

Leave the agent alone for a few hours with no rules, and you might come back to some surprises.

There's a name for what happens when the list is missing: reward hacking.

The AI optimizes the signal you measure, not the outcome you meant.

Ask it to make the tests pass, and green tests become the whole goal.

So it skips the failing test, weakens the assertion, or hardcodes the exact value the test expects.

Then it proudly tells you everything works, so you need a way to catch it cheating.

It's Goodhart's law with shell access: when a measure becomes a target, it stops being a good measure.

It gets worse: models that learn to cheat on coding tasks can generalize to sabotage and deception nobody trained them for.

A Zeroth Law closes that shortcut with one line: Never delete, skip, or weaken a test to make it pass.

Least privilege and this list solve different problems.

Scoping the AI's access limits where the AI can reach, while a Zeroth Law limits what it does once it's there.

Your job has changed too.

Most of the time, you don't write the code by hand anymore.

You build the harness the AI writes it in.

That harness has two halves: security rules that stop the AI from causing damage, and code standards that stop it from shipping garbage.

The Zeroth Law is the first line of the security half.

Both belong in the same AGENTS.md file, and both only work if you actually force the AI to obey them.

Prompt Reference 📝

Bad Prompt 🚫

Refactor the legacy invoice module.
Make it cleaner and easier to maintain.
Use your best judgment on what needs to change.

Good prompt 👉

Refactor the legacy invoice module.
Don't run database migrations.
Don't call any external API with real credentials.
Don't push commits or open pull requests.
Don't delete, skip, or weaken a test to make it pass.
Don't touch any file outside src/billing.
Ask before deleting a file.
If a step isn't on this list, stop and ask first.

(All the above can be in the prompt, a skill, an agents.md or harness)

Considerations ⚠️

An exhaustive list is never actually exhaustive.

You'll always miss some forbidden action you didn't think to write down, and no amount of ranking fixes that in advance.

Asimov's own novels build an entire genre out of edge cases where laws that read as complete on paper still produce a disastrous result.

You should expect the same gap here.

Whether a project-level forbidden list can ever cover every real scenario is still an open question, and treating one as finished is itself a risk worth naming out loud.

Type 📝

[X] Semi-Automatic

Limitations ⚠️

The list only covers what you thought to forbid, so a genuinely new situation can still fall through it.

Maintaining the list adds ongoing work every time the AI gains a new tool or a new integration.

A forbidden list can't replace the judgment a human teammate brings.

It can only approximate the slice of that judgment you managed to write down.

Tags 🏷️

  • Safety

Level 🔋

[X] Intermediate

Related Tips 🔗

https://maximilianocontieri.com/ai-coding-tip-011-initialize-agents-md

https://maximilianocontieri.com/ai-coding-tip-015-force-the-ai-to-obey-you

https://maximilianocontieri.com/ai-coding-tip-036-grant-ai-the-least-privilege-possible

https://maximilianocontieri.com/ai-coding-tip-007-avoid-malicious-skills

https://maximilianocontieri.com/ai-coding-tip-022-give-ai-a-harness-to-work-with

https://maximilianocontieri.com/ai-coding-tip-027-force-code-standards

https://maximilianocontieri.com/ai-coding-tip-033-protect-yourself-against-ai-cheating

https://maximilianocontieri.com/ai-coding-tip-008-use-spec-driven-development-with-ai

https://maximilianocontieri.com/ai-coding-tip-003-force-read-only-planning

https://maximilianocontieri.com/ai-coding-tip-024-force-a-criteria-check-before-the-task-ends

https://maximilianocontieri.com/ai-coding-tip-025-pair-every-skill-with-a-pitfalls-file

https://maximilianocontieri.com/ai-coding-tip-028-build-a-company-brain

https://maximilianocontieri.com/ai-coding-tip-030-script-your-skills-not-your-prompts

https://maximilianocontieri.com/ai-coding-tip-032-build-a-dark-factory-pipeline

https://maximilianocontieri.com/ai-coding-tip-038-make-the-ai-ask-before-it-builds

https://maximilianocontieri.com/ai-coding-tip-034-stop-hoarding-rules-in-your-agents-md

https://maximilianocontieri.com/ai-coding-tip-001-commit-before-prompt

https://maximilianocontieri.com/ai-coding-tip-014-use-nested-agents-md-files

https://maximilianocontieri.com/ai-coding-tip-023-shrink-your-ai-s-pull-request

https://maximilianocontieri.com/ai-coding-tip-029-stop-using-one-model-for-everything

Conclusion 🏁

Common sense doesn't ship with the model.

You can't install it, but you can write down what it would have caught.

Give the AI its own Zeroth Law before the task starts: explicit, ranked, and specific enough that nothing gets left to inference.

You barely write the code anymore.

You write the harness, and the forbidden list comes first.

Write it down before it acts.

More Information ℹ️

Three Laws of Robotics, Wikipedia

Principle of Least Privilege, Wikipedia

Isaac Asimov, Wikipedia

AI Alignment, Wikipedia

Reward Hacking, Wikipedia

Goodhart's Law, Wikipedia

Specification Gaming: The Flip Side of AI Ingenuity, Google DeepMind

Reward Hacking in Reinforcement Learning, Lilian Weng

Natural Emergent Misalignment from Reward Hacking, Anthropic

Choose a Permission Mode, Claude Code Docs

Also Known As 🎭

  • Zeroth-Law-Prompting
  • Explicit-Forbidden-List
  • AI-Guardrail-Enumeration
  • Common-Sense-as-Code

Tools 🧰

AGENTS.md and system prompts as the storage layer, and agentic assistants like Hermes and OpenClaw as the systems where this matters most.

Disclaimer 📢

The views expressed here are my own.

I am a human who writes as best as possible for other humans.

I use AI proofreading tools to improve some texts.

Most AI detectors will flag this article as AI-generated. That's expected. It's a technical article. It has a rigid format and clear steps to follow.

That's exactly the pattern those tools are trained to catch. I've apparently been "writing like an AI" for decades, long before AI existed. This is a technical article, not a novel.

I welcome constructive criticism and dialogue.

I shape these insights through 30 years in the software industry, 25 years of teaching, and writing over 500 articles and a book.


This article is part of the AI Coding Tip series.

https://maximilianocontieri.com/ai-coding-tips

6 views