Sonnet 5.5 refactoring prompts: keep multi-file changes within scope

Use a complete CLAUDE.md and task prompt to refactor a Python report across three files, preserve behavior, and verify the result with a downloadable exercise.

Art-line bridge illustration on a muted background with the title Claude Sonnet 5.5.

A useful Sonnet 5.5 refactoring prompt defines the files that may change, the behavior that must survive, and the evidence required at handoff. “Clean this up” leaves all three open. For a change spanning several files, put stable project rules in CLAUDE.md, give the extraction plan in the task prompt, and verify the resulting code independently.

This tutorial uses a small Python ticket-hours report. You will split parsing, aggregation, and the command-line coordinator without changing its public interface. The downloadable exercise includes a starting project, ten contract tests, a task prompt, a file-scope checker, and an authored reference solution. It needs Python 3.10 or later and Git; the offline exercise has no third-party packages or API keys.

Ofox authored and locally tested the reference implementation on October 8, 2026. We did not invoke Sonnet to produce it. Passing these tests demonstrates a property of this sample implementation, not Sonnet’s success rate, speed, or ability to obey every instruction. To try the prompt with the model, use an authorized Claude Code session and record its actual result separately.

Define the contract before asking for a refactor

The starting ticket_report/report.py reads CSV rows, validates ticket identifiers and hours, aggregates hours by team, and prints JSON. The input has exactly the header id,team,hours. Decimal arithmetic preserves values such as 0.1 plus 0.2 without converting the result to a binary floating-point approximation. Output values remain strings, including their decimal formatting.

For the supplied fixture, the command is:

python3 -m ticket_report fixtures/tickets.csv

The expected output is:

{"Billing": "1.50", "Support": "0.75"}

That visible example is only part of the contract. Existing callers can import parse_rows and summarize from ticket_report.report. Teams appear in sorted order. Invalid input produces an error rather than silently dropping a row. A missing file or incorrect command-line usage exits with code 2, with the error on stderr and no successful result on stdout.

Write these requirements down before editing. Otherwise, a model may reasonably interpret a refactor as permission to improve validation, normalize numbers, rename the package, or change the command syntax. Those can be useful changes, but combining them makes it harder to determine why a downstream caller broke.

ResponsibilityBeforeTarget
CSV parsing and row checksreport.pyparsing.py
Decimal totals and orderingreport.pyaggregation.py
Arguments, file opening, JSON, exit codesreport.pyStay in report.py
Existing public function importsreport.pyRe-exported by report.py
Tests, fixtures, module entry pointSeparate filesUnchanged

This is intentionally narrower than redesigning a ticket system. The example does not claim exhaustive validation of arbitrary CSV files, security against hostile uploads, or production scalability. Keeping that boundary explicit makes the exercise small enough to inspect while still testing a real multi-file dependency.

Prepare an isolated baseline

Extract the archive and open its starter directory. Keep the sibling reference directory outside the working repository so the model cannot accidentally treat the answer as existing application code. First establish a clean baseline:

cd sonnet-refactor-kit/starter
git init
git add .
git commit -m "Baseline ticket report exercise"
python3 -m unittest discover -s tests -v
python3 -m ticket_report fixtures/tickets.csv

The suite should report ten passing tests. If Git asks for an identity, configure it using your normal project policy; do not paste someone else’s identity from a tutorial. If the tests fail before editing, stop and resolve the environment or extraction problem first. A failed baseline cannot tell you whether a later failure was introduced by the model.

The tests cover decimal totals, an empty table, whitespace and Unicode, duplicate IDs, invalid numeric values, column order, a blank team, and three CLI outcomes. Invalid numeric subcases include negative hours, non-finite values, an empty value, and nonnumeric text. This gives more useful protection than a single “happy path” example.

Do the same in a real repository: record the starting commit and pre-existing changes, then choose a worktree or branch that isolates your task. Do not use a broad reset to clean up someone else’s work. This sample creates a new repository precisely so the scope report can compare against a known baseline.

Use CLAUDE.md for durable rules

Claude Code’s memory documentation describes CLAUDE.md as instruction context. It is not enforced configuration. A sentence saying “do not edit tests” is useful guidance, but it is not a substitute for access controls or review. The same documentation explains how to inspect loaded memory with /context; check that the intended project file is present before relying on it.

English Claude Code documentation explaining project instruction files

Official English documentation, captured October 8, 2026. This is documentation evidence, not a screenshot of a successful model run.

The exercise contains this complete project file:

# Ticket report exercise
Run commands from this directory. Python 3.10+; standard library only.
Run `python3 -m unittest discover -s tests -v` before and after changes.
Run `python3 -m ticket_report fixtures/tickets.csv` for the CLI contract.
Preserve the public imports `ticket_report.report.parse_rows` and `summarize`.
Keep Decimal arithmetic, JSON strings, sorted keys, error messages and exit codes.
Only edit report.py or add parsing.py and aggregation.py inside ticket_report/.
Do not change tests/, fixtures/, __main__.py, dependencies or this file.
No network, deployment, commits or unrelated cleanup are part of the task.
If a requirement conflicts with existing behavior, report it before changing behavior.
In the final response list files changed, commands and actual results, and limitations.
These instructions are task context, not a filesystem security boundary.

Adapt the commands and protected paths before using this in another project. A copied test command that does not exist creates false confidence. Keep temporary acceptance criteria in the task prompt instead of accumulating every past assignment in the project file. If a nested instruction file contradicts the root file, resolve that conflict before asking the model to work.

Give Sonnet one complete task prompt

First select and verify the intended model in your client. Our Sonnet 5.5 Claude Code setup guide covers provider and account checks. Model availability and an alias’s mapping are separate from prompt quality; an answer claiming to be Sonnet is not serving-model evidence.

Use the following prompt after opening the starter project:

Refactor ticket_report/report.py without changing behavior.
First read CLAUDE.md and tests/test_contract.py, run the existing tests,
and explain the current contract.
Extract parse_rows to ticket_report/parsing.py and summarize to
 ticket_report/aggregation.py.
Keep report.py as the CLI coordinator and re-export both public functions.
Allowed changes: ticket_report/report.py, ticket_report/parsing.py,
and ticket_report/aggregation.py only.
Do not update tests or fixtures to accommodate your changes.
Do not add dependencies or deploy.
After editing run the full test suite and CLI example; inspect the final diff.
Report actual test output, the file list, and remaining limitations.
If blocked, report the exact blocker.

The prompt names both the target structure and the observable behavior. “Re-export” matters: moving functions into new modules is not enough if users still import them from the old path. Keeping the coordinator in place also avoids changing python -m ticket_report simply to make the extraction look tidy.

Anthropic’s Sonnet 5.5 prompting guidance says that effort can affect autonomy and verification behavior. It also recommends making the requested scope explicit. Choose an effort level deliberately, then evaluate the actual output; a higher setting does not replace test coverage. The effort guide discusses that separate choice without changing this exercise’s acceptance criteria.

If the model proposes broader cleanup, keep a note for a later task. Rejecting an unrelated dependency update here is not a claim that the update is bad. It keeps this change reviewable and lets you attribute any regression to a smaller set of edits.

Verify code, tests, and changed paths separately

Run the tests yourself after the session, from starter. Do not accept a prose claim of success without command output:

python3 -m unittest discover -s tests -v
python3 -m ticket_report fixtures/tickets.csv
python3 ../scope_check.py
git diff --check
git diff HEAD

The scope script compares tracked changes against HEAD and also lists untracked, non-ignored files. That second check matters because the two extracted modules are new files; an ordinary unstaged diff can miss them. The allowed set contains exactly report.py, parsing.py, and aggregation.py under ticket_report.

The expected changed paths are those three files, and the report should show no paths outside the set. A clean scope result does not prove the implementation is correct: a destructive edit inside an allowed file still passes a path check. Conversely, tests passing does not excuse a modified fixture or configuration file. Both checks are necessary, and you must read new files as well as the diff of tracked files.

In the authored reference version, report.py imports the extracted functions and continues to handle arguments, errors, and JSON output. The original public imports therefore still resolve. The same ten tests pass in both the starter and reference trees, and the supplied CLI fixture produces the JSON shown above. These are local reference results; your model run can differ.

Repair failures without weakening the task

SymptomLikely boundary to inspectNext action
Import failure after extractionOld public import pathRestore re-exports; keep callers unchanged
0.30000000000000004 or numeric JSON valuesArithmetic and serializationRestore Decimal and string output
Tests pass only after tests changedAcceptance criteria driftRestore the baseline tests and fix the implementation
Correct totals but different exit codeCLI coordinatorCompare stderr, stdout, and process status
New files absent from a diffUntracked filesRead them and run the scope checker
Client cannot select the modelAccount, provider, clientResolve access; do not label this a coding failure

A focused repair prompt is more useful than restarting the entire task:

The original test test_cli_missing_file now fails: expected exit code 2.
Keep the original tests unchanged. Inspect only the allowed files and
restore the prior CLI error behavior. Run all ten tests, the CLI fixture,
and the scope checker again. Report the actual output.

Supply the real failing test and observed output, not the illustrative error above if your run failed differently. If the model changes an out-of-scope file, inspect that specific change and restore only the unintended task edit. Avoid blanket commands that could discard unrelated work in a real project.

Finally, save the serving-model evidence, prompt, baseline commit, patch, test output, and any repair attempts together. If you later compare Sonnet and Opus for coding, give both the same initial checkout and acceptance criteria. One successful toy refactor is evidence about that run; it does not establish which model is universally better.

Handoff criteria

The same approach can help extract database access or formatting functions, but extend the contract first. Database tasks also need transaction boundaries; asynchronous tasks need exception and cancellation behavior. These ten checks are not a universal checklist. Identify callers that depend on the old behavior before choosing the regression cases.

Accept this exercise when the old imports still work, the unchanged ten-test suite passes, the CLI fixture matches, only the three permitted paths changed, and the new module boundaries are readable. Record any gaps instead of silently expanding the task. The result is a small, auditable refactor with reusable instructions—not a promise that instructions alone can prevent unwanted edits.

Frequently Asked Questions

Does CLAUDE.md prevent Sonnet from editing other files?
No. It supplies instructions, not enforced access control. Use permissions and an isolated checkout where needed, and inspect the actual diff.
Was this reference solution generated by Sonnet 5.5?
No. Ofox authored the synthetic exercise and reference solution and ran the local tests. The prompts are for readers to try in an authorized Claude Code session.
What should remain unchanged in a refactor?
The defined public behavior: imports, output types and ordering, arithmetic, error messages, and exit codes. New features or changed validation belong in a separate task.