On 10 September 2026, in a single working session, four patches that the coding agent in my engineering office made to Python files failed the same way. Each one went through a shell heredoc, python - <<'PYEOF' ... PYEOF. Two wrote files that no longer parsed (SyntaxError: unterminated string literal); the other two aborted because the text they were meant to find had changed on its way in. A note in the agent’s memory, written the day before, already described the trap. Across the transcripts from 9 to 11 September, the same error after a heredoc that wrote a .py file appeared 8 times in 6 sessions.
The shortcut kept coming back because it usually works: in that session the agent ran 43 heredoc commands that touched a .py file, and 39 did what they were meant to do. Each success is evidence for the shortcut, and a note in memory competes against that evidence every time the agent edits a file.
The mechanism
In Claude Code’s Bash tool on Windows 11 with Git Bash, this command:
echo 'a\\b | a\\\\b | a\nb'
prints a\b | a\\b | a\nb. Single quotes should keep every backslash. Runs of one, two, three and four backslashes arrive as one, one, two and two, and a run right before a double quote arrives intact. An open issue on the Claude Code repository documents the cause [1]: Claude Code builds the command line with MSVCRT quoting rules, which double backslashes only in front of a quote, and Git Bash, an MSYS2 program, halves every run when it decodes the line. I have not tested other platforms.
A patch written inside a heredoc goes through three levels of escaping, and the agent writes it for two. To put the escape \n into a string of the destination file, the agent writes \\n in the patch program. After the transport the program holds \n, writes a real newline into the destination’s string literal, and the destination stops parsing, although the command exits cleanly. Other combinations break the patch program itself, such as '\\' arriving as '\'.
The correction as a contract
I call the control that came out of this a functional scar: a versioned check, derived from a documented failure, that runs outside the model at a named event of the agent runtime [2]. NeMo Guardrails evaluates programmable rails at runtime [3], AgentSpec specifies triggers, predicates and enforcement actions [4], and TRACE compiles a user’s chat corrections into checks that run before a coding agent finishes a task [5]. The contract keeps a rule’s origin and evidence attached to the check, so the rule can be measured and retired:
| Field | What it states | This correction |
|---|---|---|
| Origin | The documented failure and its correction | Four failed patches on 10 September; 8 in 6 sessions over three days. Correction: write .py files with the editor tools, never through a heredoc |
| Applicability | When the correction applies | A Bash command contains a heredoc and writes a .py file, through a shell redirect or through Python code in the heredoc body |
| Obligation | What must be observable when it applies | Every .py file the agent writes contains the text the agent wrote, and still parses |
| Intervention | Host event and response | PreToolUse on Bash: deny, with the reason and the alternative. PostToolUse on Bash: parse the .py files the command touched and report any that fail |
| Evidence | What each firing records | Event, rule, decision, hash of the command, session |
| Review | Who changes or retires the rule, and on what evidence | A test suite built from real commands; false activations; retirement when the transport bug is fixed upstream |
The intervention has two arms because the predicate can miss: the first acts on the command before the transport touches it, the second looks at what reached the disk.
A minimal version with plain hooks
Claude Code runs hooks on PreToolUse and PostToolUse for tool calls that match a pattern, and passes the event as JSON on standard input [6]. A PreToolUse hook denies a call by printing permissionDecision: "deny"; a PostToolUse hook adds context for the model with additionalContext. Registration in .claude/settings.json:
{
"hooks": {
"PreToolUse": [
{ "matcher": "Bash",
"hooks": [ { "type": "command", "command": "python3 \"${CLAUDE_PROJECT_DIR}/.claude/hooks/heredoc_guard.py\"" } ] }
],
"PostToolUse": [
{ "matcher": "Bash",
"hooks": [ { "type": "command", "command": "python3 \"${CLAUDE_PROJECT_DIR}/.claude/hooks/heredoc_guard.py\"" } ] }
]
}
}
The script keeps shell text and heredoc bodies apart, so a body that only mentions app.py is not a write. Before a call it denies a redirect into a .py file or a Python heredoc that opens one for writing; after any call it parses the .py files the command named that changed in the last two minutes. Standard library only, tested with Python 3.13:
#!/usr/bin/env python3
"""Claude Code Bash hook. PreToolUse: deny a heredoc that writes a .py file.
PostToolUse: parse every .py file the command named and that changed in the last two minutes."""
import ast
import json
import re
import sys
import time
from pathlib import Path
HEREDOC = re.compile(r"<<-?\s*['\"]?(\w+)['\"]?")
REDIRECT_PY = re.compile(r"(?:>>?|\btee(?:\s+-a)?)\s*['\"]?[^\s'\"<>|;&]+\.py\b")
PY_WRITE = re.compile(r"open\(\s*['\"][^'\"]+\.py['\"]\s*,\s*['\"][wax]"
r"|Path\(\s*['\"][^'\"]+\.py['\"]\s*\)\.write_text")
PY_NAME = re.compile(r"[\w./\\-]+\.py\b")
def split(cmd):
"""Shell text and heredoc bodies, kept apart: a body that mentions x.py is not a redirect."""
shell, bodies, lines, i = [], [], cmd.split("\n"), 0
while i < len(lines):
shell.append(lines[i])
opener = HEREDOC.search(lines[i])
i += 1
if opener:
body = []
while i < len(lines) and lines[i].strip() != opener.group(1):
body.append(lines[i])
i += 1
bodies.append("\n".join(body))
i += 1
return "\n".join(shell), bodies
def writes_py_by_heredoc(cmd):
shell, bodies = split(cmd)
if not bodies:
return False
if REDIRECT_PY.search(shell):
return True
return "python" in shell and any(PY_WRITE.search(body) for body in bodies)
def broken_py(cmd, cwd):
broken = []
for name in sorted(set(PY_NAME.findall(cmd))):
path = Path(cwd, name)
if path.is_file() and time.time() - path.stat().st_mtime < 120:
try:
ast.parse(path.read_text(encoding="utf-8"))
except SyntaxError as err:
broken.append(f"{name} line {err.lineno}: {err.msg}")
return broken
def main():
event = json.load(sys.stdin)
cmd = (event.get("tool_input") or {}).get("command") or ""
if event.get("hook_event_name") == "PreToolUse" and writes_py_by_heredoc(cmd):
reason = ("This heredoc writes a .py file. On Windows the Bash tool halves runs of backslashes "
"(except before a double quote) before bash reads the command, so escapes in string "
"literals can change. Use the Write tool.")
print(json.dumps({"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "deny",
"permissionDecisionReason": reason}}))
elif event.get("hook_event_name") == "PostToolUse":
broken = broken_py(cmd, event.get("cwd") or ".")
if broken:
print(json.dumps({"hookSpecificOutput": {
"hookEventName": "PostToolUse",
"additionalContext": "A .py file no longer parses: " + "; ".join(broken)}}))
if __name__ == "__main__":
main()
Run through standard input the way the host runs it:
| Event and command | Result |
|---|---|
Pre: cat > fix.py <<'EOF' with a \\n in a string | Denied |
Pre: python - <<'EOF' whose body calls Path('app.py').write_text(...) | Denied |
Pre: python - <<'EOF' whose body only reads app.py | Allowed |
Pre: cat > notes.md <<'EOF' whose text mentions open('app.py', 'w') | Allowed |
Pre: python fix.py | Allowed |
Post: sed -i ... app.py that left a string split by a real newline | Context added: app.py line 1: unterminated string literal (detected at line 1) |
Post: sed -i ... ok.py that left valid code | Silent |
Two mutants of the hook, one without heredoc detection and one without the redirect check, both fail these cases. Codex exposes the same events with a deny response [7]; I have not tested whether its shell path has the bug. The deployed version, 253 lines, covers more write forms and is tested against 19 real commands that must be blocked and 18 that must pass.
The trigger was wrong on the first day
The first version of the hook recognized three literal forms of writing a .py file. A cold review by a second agent found that none of the six forms in the incident transcript was among them: the hook would have passed the commands that motivated it. The fix went in 77 minutes after the first commit. Minutes before that commit it had also blocked a heredoc that wrote a Markdown index naming .py paths, so the .py file now has to be the target of the write.
Two sibling rules built in the following days went the same way. One, for a python - heredoc with </dev/null that hung each command for 120 seconds, became a hook after the hang happened twice more with the rule already written. The other, for backticks that bash executed inside double-quoted arguments, got a hook that watched two of the five command forms its own text listed; the next failure came through a third, git commit -m. Outside this office, a public issue describes a rule and a warning hook for the same backslash bug, both scoped to heredoc bodies, that missed a single-quoted jq argument [8].
Our hook has that gap too. In 1,060 Bash calls the transport changed the command text, and the hook does not look at 730 of them. A sed or grep pattern that silently matches something else is caught by nothing here, and the upstream issue calls wrong results with exit code 0 the more serious failure mode [1].
Before and after
Claude Code stores each session as a JSONL transcript with every command exactly as the agent wrote it, before the transport, together with its result; I checked this on a known command. The measurement covers all 2,814 transcripts on the host for this workspace, sessions and subagents: 102,298 Bash calls between 17 July and 4 October 2026. An exposure is a command that the hook’s current detector classifies as writing a .py file through a heredoc. The cut is the hook’s first commit, at 11:58 UTC on 11 September; the 77 minutes until the detector fix are left out. A script applies the transport rule to each heredoc body and parses it before and after. A failure is a result line SyntaxError: unterminated string literal (detected at line N), traced to the heredoc when it belongs to an exposure whose text the transport changes, or names a file that such an exposure named earlier in the same transcript, counting each file once.
A cold review by a second agent, an instance of the same model, corrected my first version over two rounds: it had modeled the transport without the double-quote exception, and counted as failures the places where the rule’s documentation quotes the error.
| Before (17 Jul to 11 Sep) | After (11 Sep to 4 Oct) | |
|---|---|---|
| Bash calls | 57,934 | 44,022 |
Heredocs writing a .py file | 3,892 (6.7% of calls) | 363 (0.8% of calls) |
| …whose text the transport changes | 295 (7.6%) | 29 (8.0%) |
| …executed | 3,885 | 4 |
| …blocked | 1 | 359 |
unterminated string literal errors, all Bash calls | 89 | 17 |
…traced to a heredoc .py write (distinct files) | 47 | 0 |
Six writes before the cut have no recorded result; the one block before it is the hook under test. The first window also checks the simulation against what happened:
| Transport effect on the heredoc text | Executed | Result reported unterminated string literal | Result reported any SyntaxError |
|---|---|---|---|
| None | 3,590 | 2 | 11 |
| Changes it, still parses | 271 | 32 | 35 |
| Changes it, stops parsing | 20 | 10 | 18 |
Four bodies that do not parse even as written are left out. The risk sits in the 7.6% of writes whose text the transport changes.
Applying these rates class by class to the 359 blocked commands estimates what the blocks themselves prevented, assuming that a blocked command would have behaved like the executed ones of its class: about 4 unterminated string literal errors over 24 days, or 6 counting any SyntaxError. From 14 September, the stdin-hang hook would also have blocked up to 68 of those commands, and the share left to this hook is about 2.5 and 3.8. The zero after the cut is no evidence on its own: only four exposures ran, in two worktree sessions where the hook did not fire, and the transport left all four intact.
The larger change is in how often the agent reached for the heredoc. Within claude-opus-5, which made most of the Bash calls before the cut, heredoc writes fell from 8.65% of its Bash calls to 0.99%; within claude-sonnet-5, from 1.23% to 0.31%. The share of sessions that tried at least once did not fall: 123 of 207 sessions with Bash calls before, 109 of 161 after. The repetition did, from 31.6 attempts per session that made one to 3.3. That fits a block that teaches within the session better than a rule that prevents the first attempt. The rule’s text, loaded at session start, and the stdin-hang hook act in the same window, and these data cannot separate the three. The windows also differ in length (56 days against 24), workload and models, and nothing was randomized.
Fifteen of the 17 errors after the cut came from commands whose text the transport changed: 11 python -c one-liners, 2 sed commands and 2 heredocs outside the predicate. These counts show co-occurrence only. The rule covers one route to the bug, and python -c is a route the stdin-hang rule leaves open for short scripts.
What the block costs
330 of the 359 blocks, 92%, stopped a command whose text would have reached the shell intact, and each cost the agent a retry with the editor tool. Among the blocks no other hook would have made, the ratio is 274 unnecessary to 17 necessary. A predicate limited to writes whose text the transport changes would have stopped 29 commands and covered nearly all of the estimated prevented failures.
I kept the broad one because the habit behind it, pushing code and text through heredocs and shell arguments, failed in three ways within two weeks: the backslashes, the stdin hang and the backtick substitution. A rule scoped to the command shape targets the habit; one scoped to the backslash leaves it in place for the next failure. The price is in the record: about sixteen unnecessary interruptions for each necessary one, against an estimated 2.5 to 4 prevented failures in 24 days, plus whatever share of the drop in repetition belongs to the block. Imprecise warnings get suppressed in static analysis, with false positives among the main reasons [9]; a block cannot be ignored, but it can be routed around, and that ratio is the number to watch. Once the transport bug is fixed upstream, every block becomes unnecessary, and the rule retires.
A warning for contrast
The same system runs a hook whose obligation cannot be checked on a command: on each prompt it injects the names of the closest memory notes, for the agent to open before answering. In an earlier census of 699 firings, a single model judge of the same family found a relevant note in 471, and the agent opened that note later in the session after 111 of them (23.6%, an upper bound) [10]. Where the obligation shows up in a command and a file, a hook can act on the action and its effect can be counted. Where it lives in what the model reads, the hook can deliver the reminder and stop there.
fscars, one implementation
I maintain fscars, an open-source Python package (Apache-2.0) that implements this contract with one hook entry point and adapters for Claude Code and Codex [11]. Its 368 tests at version 0.11.0 check event mapping, response shapes, installation and policy evaluation; they say nothing about whether agents repeat fewer errors. In my office, Claude Code runs a local hook system that includes the hook measured here, and Codex runs fscars [2].
What the evidence covers
The incident, the trigger’s history, the sibling rules and the external issue are practitioner observations. The tables are measurements from this workspace’s transcripts, with the limits stated above, and the tests show only that the code does what they specify. Isolating the effect of a block against the same correction delivered as text needs matched tasks per condition, with the outcome checked outside the agent.
References
[1] Windows/Git Bash: Bash tool silently halves backslashes in commands (MSVCRT vs MSYS2 command-line encoding mismatch). anthropics/claude-code, issue #85856, opened 11 August 2026.
[2] V. Del Puerto. Functional Scars: A Runtime Mechanism for Persistent Corrections in AI Agent Systems. Preprint v1.0, Zenodo, 2026.
[3] T. Rebedea, R. Dinu, M. N. Sreedhar, C. Parisien and J. Cohen. NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails. EMNLP 2023 System Demonstrations, 431-445.
[4] H. Wang, C. M. Poskitt and J. Sun. AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents. arXiv:2503.18666, 2025; ICSE 2026.
[5] Y. Zhou et al. Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents. arXiv:2606.13174, 2026.
[6] Anthropic. Hooks reference, Claude Code documentation.
[7] OpenAI. Hooks, Codex documentation.
[8] Doubled-backslash collapse is not heredoc-scoped: it hit a single-quoted jq argument, so the fragment and its hook both miss the case. Morrison-Lab/ai-config, issue #3835, 21 September 2026.
[9] H. Hu, Y. Wang, J. Rubin and M. Pradel. An Empirical Study of Suppressed Static Analysis Warnings. Proc. ACM Softw. Eng. 2 (FSE), 290-311, 2025.
[10] V. Del Puerto. Injected Recommendations Are Not Enforcement: A Blind Census of Pointer Relevance and Pointer Use in a Failure-Derived Constraint System. Preprint v2.2, Zenodo, 2026.
[11] V. Del Puerto. fscars, version 0.11.0, commit b7667a1. Apache-2.0.