AI Automation · 5 min read

Prompt, context, harness, loop: the four disciplines behind AI that actually works

By Chad ·

The plateau nobody warns you about

The first month is great. Something that took an hour takes ten minutes, and it feels like the start of a curve.

Then it flattens, and what you start hearing around the shop is some version of these four complaints:

Those four sound like one problem with one fix — get better at prompting. They aren't. They're four different problems, and three don't respond to prompting at all. That's why the plateau feels so stubborn: most people apply the rung-one fix to a rung-three problem and conclude the technology is overhyped.

Four problems wearing one costume

Each complaint belongs to a different discipline, and each has a name: prompt engineering, context engineering, harness engineering, loop engineering. The names sound like they belong to a software team. They don't — each describes a decision an owner makes about how the work gets done.

They stack, and the order matters. To keep this out of the abstract, one job runs through all four: a quote for the Neely job, four times, one rung apart.

Rung one — prompt engineering: how you ask

This is the one everybody knows: the quality of a single request. What you're asking for, in what format, with what constraints, for whom.

You paste in the job — three-ton condenser, 1970s house, air handler in the attic — and ask for a quote. What comes back is competent and entirely forgettable. Three tiers, filler warranty language, a start date of "at your earliest convenience." Nothing is wrong with it. Nothing in it is yours.

Prompting is real, it's learnable, and it's worth an afternoon. It's also the only rung you can climb without changing anything about your business — which is exactly why most people never leave it, and why the payoff caps out fast. A better-worded request buys one better answer, and the next request starts from zero again.

That ceiling is low, and it sits close enough to where most people stop that they mistake it for the ceiling of the technology. It isn't. It's the ceiling of rung one.

We wrote the practical version already: why your prompts feel generic, and how to fix it.

Rung two — context engineering: what it knows before you ask

The second complaint isn't about the request. It's about everything the request assumes.

Give the same tool standing context — your price floor, your warranty wording, the rule that nothing gets discounted past 15% without a flag — and the Neely quote comes back sounding like you wrote it. Not because you asked better, but because it wasn't starting from nothing.

Context is the baseline: what the system knows before anyone types a word. That's the difference between a good answer and every answer starting good, and it's where the work becomes an asset you own.

The how: standing context beats better prompts, and the deeper version — teach AI your playbook.

Rung three — harness engineering: what it's allowed to touch

Third complaint, and the first one that prompting cannot touch at all: it can tell you exactly what to do and it cannot do any of it.

The harness is the wiring around the AI — what it can read, what it can change, and where it has to stop. On the Neely quote, it pulls today's condenser cost from your supplier's price list instead of guessing, and drops the quote into your CRM as a draft. It also cannot send it. That limit isn't a shortcoming; it's the design.

This is where AI stops producing advice and starts producing work. What your AI is allowed to touch.

Rung four — loop engineering: where corrections go

The last complaint tells you the whole thing has stalled: it makes the same mistake every week.

On the Neely quote it's the after-hours rate, stated wrong every time. Fixing it in the draft takes twenty seconds and you've done it nine times. Fixing it in the playbook takes two minutes, once.

That's the discipline. Note what this rung acts on: rung two is the baseline as of the day you wrote it, and the loop is the only thing that legitimately edits that baseline. Without one, your standing context is permanently a snapshot of March. Fix it once.

Why it's a ladder and not a list

Each rung makes the one below it permanent:

That last line is the argument. Rungs one through three get you a system that performs at a level; rung four is the only one that changes the level.

Which rung are you on

Whichever complaint is loudest is where you're stuck. That's the diagnosis, and it's usually obvious the moment you stop assuming everything is a prompting problem.

Don't skip rungs. A harness built on no context automates generic output faster and with more confidence — worse than the manual version, because now it's wrong at scale and nobody's reading closely. And a loop with nowhere to write corrections back to is just you, complaining on a schedule.

This week

  1. Say your loudest complaint out loud, in the words above.
  2. Read the one post that answers it. Only that one.
  3. Do what it says on one real task — not five.

Then check whether the complaint moved down a rung. That's the tell.


If you'd rather not work out which rung you're stuck on by yourself, that's a 30-minute conversation. Book a call and bring the task that's been annoying you longest.