AI Refactoring: Breaking Up a God Class
Every codebase I've inherited from another dev has had one: a 1,200-line OrderService or UserManager class that handles validation, database writes, email sending, and three unrelated side effects, all in the same file. AI refactoring tools promise to break these apart for you, and lately I've been using Claude and GPT-4 to do exactly that on a client project — a Laravel order processor that had grown into a genuine god class over two years of "just add it here for now." It works, but only if you stop treating the AI like a magic "clean this up" button.
Why God Classes Resist a Quick Fix
A god class is hard to refactor by hand because the coupling is invisible until you try to move something. Pull out the email logic and you discover it silently depends on a $this->status flag three methods away. Pull out the validation and half of it turns out to double as business logic nobody documented. This is exactly the kind of problem AI is decent at — it can read the whole file in one pass and spot the hidden dependencies faster than you can scroll through 1,200 lines — but it's also exactly the kind of problem where a bad prompt produces a refactor that looks clean and breaks production.
I learned this the expensive way on that Laravel project. My first pass at AI-assisted refactoring shipped a "cleaned up" OrderService that passed every existing test and still managed to stop sending confirmation emails, because the AI moved the email call into a method that only fired on the happy path of a switch statement it didn't fully understand.
The tests passing is the part that should worry you more than the bug itself. A god class that's been patched for two years almost never has test coverage for its edge cases — the tests cover the paths someone actually thought to write tests for, which is rarely the tangled switch branch that's been quietly broken-but-working since 2024. An AI refactor that only has those tests as a safety net will happily "improve" code whose real behavior nobody wrote down anywhere.
The Prompt That Produces Garbage
Here's roughly what I asked the first time, and what it produced:
Prompt: "Refactor this OrderService class to follow single
responsibility principle. Split it into smaller classes."
The model happily obliged — it extracted an OrderValidator, an OrderNotifier, and an OrderRepository, and every one of them looked plausible in isolation. What it didn't do: preserve the exact order of side effects, which mattered because the notification depended on a database write that hadn't committed yet in the new version. A vague "split this up" prompt gives the AI permission to reorganize execution order, and it will, because nothing in the prompt told it that order was load-bearing.
The Prompt That Actually Works
The fix isn't a smarter model, it's a more constrained prompt — one that treats the AI as a very fast pattern-matcher that needs explicit guardrails, not a senior engineer who already knows your invariants:
Prompt: "Here is OrderService.php [paste full file]. Identify
every distinct responsibility (validation, persistence,
notification, external API calls) and list them with the exact
line ranges for each, WITHOUT changing any code yet.
For each responsibility, tell me:
1. What other responsibilities it reads state from
2. Whether its execution order relative to the others matters
and why
Do not propose the extracted classes yet — just the map."
That two-step approach — map first, extract second — is the actual unlock. Once the model lists out "notification reads $order->id, which only exists after the persistence step commits," you can see the ordering dependency before any code moves, instead of discovering it in a bug report. Only after I review that map do I ask for the actual extracted classes, one at a time, with a instruction to keep the original call order intact unless I explicitly approve a change.
I'd skip any workflow that asks an AI to refactor and reorder a god class in a single step — the two are different problems, and collapsing them is where most "AI broke my refactor" stories come from. A senior engineer wouldn't do a full extract-and-reorder in one uninterrupted pass on code they don't own the history of, and an AI shouldn't either.
Where This Still Needs a Human
The mapping step is where AI earns its keep — it's genuinely faster than manually tracing coupling across a 1,200-line file. What it still can't do reliably: know which side effects are actually required by the business versus which ones are legacy cruft nobody remembers approving. On that same project, the god class had a method that wrote to an audit log table nobody queried anymore. The AI preserved it faithfully across three refactor passes because nothing told it the table was dead. I only caught it because I happened to check the table's row-insert timestamp against its last SELECT in the query logs. That's a judgment call an AI has no way to make from the code alone, and it's exactly the kind of thing a human still has to verify before signing off on an AI-assisted refactor.
Once the extraction itself is underway, I also make the AI write a short "behavior diff" in plain English before I look at the actual diff — a list of anything that changed about when something runs, not just where the code lives. Nine times out of ten it says "no execution order changes." The tenth time is the one that saves you a production incident, because it'll flag something like "notification now fires before the database transaction commits in the failure branch" — which is exactly the class of bug that slips past a normal code review, since the diff looks like a clean move, not a behavior change.
Conclusion
Map dependencies before you ask for extraction, keep execution order explicit until you've verified it doesn't matter, and never trust a single-pass "clean this up" refactor on a class you didn't write yourself. The AI is fast at finding structure; it's still on you to know which parts of that structure are load-bearing.
0 Comments
No comments yet — be the first to share your thoughts.