Back to articles
AIEngineering

The Year Coding Agents Stopped Being a Novelty

AI went from autocomplete to agents that open PRs, run tests, and fix their own failures. Here's how I actually fit them into a real workflow without losing the plot — or the quality.

Yash Thakur
4 min read
The Year Coding Agents Stopped Being a Novelty

The short answer: coding agents now take whole bounded, verifiable tasks — branch, write, test, iterate, open a PR — so the job shifts to writing tight specs and treating every output as a proposal you review, never an automatic merge. Hand agents mechanical refactors, tests, and the well-understood 70%; keep architecture and trade-offs for yourself.

A year ago AI coding tools were fancy autocomplete. In 2026 they're agents — they read a ticket, branch, write code, run the test suite, read the failures, and iterate until it's green, then open a PR. That's a different thing entirely, and it's forced me to rethink what my job even is day to day. Here's where I've landed after actually running agents on real work, not demos.

The unit of delegation changed

With autocomplete, I delegated keystrokes. With agents, I delegate tasks — but only a specific kind: bounded, well-specified, verifiable ones. "Add a rate limit to the upload endpoint, here's the pattern we use elsewhere, make the existing tests pass and add one for the limit" is a great agent task. "Make the checkout flow better" is not — it has no definition of done, and an agent with no definition of done will confidently produce something plausible and wrong.

The skill that matters now is writing the spec. A tight task description with acceptance criteria, the files involved, and the constraints is worth more than any prompt-engineering trick. It's the same skill as writing a good ticket for a junior engineer — which, not coincidentally, is what an agent most resembles.

Verification is the whole game

When an agent can generate a hundred lines in thirty seconds, your bottleneck moves entirely to trusting those lines. I've built a hard rule: an agent's output is a proposal, never a merge. It runs the tests, I read the diff. If the diff is large and I can't hold it in my head, that's a signal the task was too big, not that I should skim it.

The failure mode I watch for on my team is "LGTM-by-exhaustion" — reviewing so much agent output that attention frays and a subtly-wrong change slips through. The guardrail is the same as it's always been: good tests, small diffs, and the discipline to actually read them.

Where agents genuinely shine

  • Mechanical refactors across many files — rename a concept, migrate an API, update a pattern everywhere. Tedious for a human, perfect for an agent, and easy to verify by diff + green tests.
  • Writing tests for existing code — point it at a module and it drafts the edge cases, which I then prune and correct.
  • The first 70% of a well-understood feature — scaffolding, wiring, the obvious parts — leaving me the 30% that needs judgment.
  • Investigation — "where does this value get set, trace it" across an unfamiliar codebase.

Where I keep the keys

Architecture, trade-offs, and anything touching why the system is the way it is. An agent optimizes for the local task; it has no stake in the coherence of the whole. It'll happily add the fifth slightly-different way of doing data fetching because each one, locally, works. Keeping the codebase a coherent thing rather than a statistical average of the internet is now an active, human job — arguably the main one.

What it means for the team

The thing I tell the engineers I work with: agents raise the floor on output and raise the ceiling on what judgment is worth. The people who thrive treat agents as leverage on the parts of the job that were never the point — the typing, the boilerplate, the grind — and reinvest the time in the parts that always were: deciding what to build, why, and whether what came back is actually right. The ones who struggle outsourced the thinking along with the typing.

Coding agents stopped being a novelty this year. They didn't replace the work; they relocated it — from writing lines to specifying intent and verifying results. That's a better job, honestly, as long as you stay the one holding the standard.

Written by Yash Thakur

Senior React Developer · 8+ years building for the web

More articles

Keep reading