There’s an agent you built a few weeks ago that you haven’t opened since.
Maybe it summarised your emails. Maybe it wrote social posts. And it worked… you remember it working, which is the part that makes it stick in your head.
Then a week got busy and you never went back.
That’s the normal outcome and almost nobody says so out loud. Every guide ends at the moment the thing runs successfully for the first time, which is a bit like a recipe that finishes when the oven door closes.
Here’s what I think actually went wrong, and it’s neither the model nor you.
You picked the tool first.
Every guide tells you to. Choose a platform, take the tour, build the demo task. So the job you end up choosing is whichever one shows the platform off best… and that is almost never a job you needed doing on a Tuesday.
The one build I still use months later is the dullest thing I own. No clever reasoning, no chain of tools, nothing you’d screenshot. The gap between that and the ones I abandoned had nothing to do with the model, the platform, or how the prompt was written.
So here’s the order I’d build in now, which is roughly the reverse of how most people start.
The short answer
To build your first AI agent, pick the job before you pick the tool. A job worth automating repeats on a schedule you can name, has a clear definition of done, and is something you’ve already done by hand enough times to recognise a bad result. Then give it the context it needs, give it a trigger that isn’t you remembering, and decide where you stay in the loop.
The platform is the last decision, and by then it barely matters which one you choose.
1. Why the tool is the last decision
Open any guide on this and step one is choose your platform. Then there’s a tour of the interface, a demo task, and a finished thing that looks genuinely impressive.
I’ve followed several of those guides. The builds worked every time.
Almost none of them survived contact with a real week.
Here’s why, and it took me a while to see it. When you start with the tool, you end up choosing a job that shows the tool off. Something with a satisfying output, ideally visual, ideally something you can watch happen. That’s a demo, and demos are chosen for how they look rather than for whether anyone needed them.
When you start with the job, you pick something dull that was already annoying you… and then the tool question answers itself, because most of them do the dull thing fine.
Anything you’re going to use every week is boring by definition. That’s not a compromise, it’s the actual signal.
2. Step one: pick a job that repeats
Here’s the test I use now. All four have to be true, and the failure is almost always the same one.
It repeats on a schedule you can name
Weekly. Every time I publish. First of the month. If you can’t say when it happens, it isn’t a job yet… it’s a topic.
“Help me with marketing” fails this instantly. “Check every Monday whether last week’s post got promoted” passes.
It has a definition of done
You need to be able to say what finished looks like before it starts.
An assistant that doesn’t know what done means will keep going past it… confidently, in whatever direction it was already heading.
You’ve done it by hand enough times to know what bad looks like
This is the one people skip, and it’s the one that decides whether you’ll trust the output.
If you’ve never done the job yourself, you can’t tell a good result from a plausible one. You’ll either accept everything, which is how bad work reaches your audience… or you’ll check everything, which costs more than doing it yourself did.
It’s near money or near the list
Not a hard rule, but a good tiebreaker when two jobs both qualify.
Prefer the one that touches an actual customer, an actual sale, or the email list. The list is the most portable thing you own, and a job that protects or grows it earns its place faster than one that tidies your files.
And four that will waste your weekend
The inverse of the tests, because these are the ideas that feel best and fail hardest.
Not the most impressive thing you can think of. Impressive jobs are usually one-offs, and one-offs don’t need agents… they need one good session.
Not something you’ve never done manually. You won’t be able to judge the output, and unjudged output is how bad work reaches your audience with your name on it.
Not something with no clear finish. “Monitor my competitors” runs forever and produces a pile. “Every Monday, tell me if any of these five sites published something about AI search” finishes.
And not a job you don’t actually have. The most common first agent is built for a workflow the person imagined rather than one they run. If you had to invent the use case, you’ll never open it twice.
Six that pass, if you can’t think of one
The tests are easy to agree with and harder to apply to your own week, so here are real ones. Each already passes all four.
| The job | Runs | Done looks like |
|---|---|---|
| Did last week’s work actually get promoted? | Weekly | Every post from last week has a status against each channel |
| Who joined the list and never got the delivery email? | Weekly | Every new subscriber is confirmed delivered, or flagged |
| Which replies did I never answer? | Friday | Every inbound message from the week is answered or listed |
| Which invoices are past due? | Weekly | Every unpaid invoice over 14 days, with client and amount |
| Which links on the site are now dead? | Monthly | Every link in the last 20 posts returns a status |
| Which drafts have gone stale? | Weekly | Every draft untouched for 30 days is listed |
If two of those apply to you, take the second one. A subscriber who joined and never received what you promised is a broken funnel losing you money silently, and it is the single most common thing I see wrong in a one-person business.
Run the four tests against whatever you were about to build instead. Most first-agent ideas die at the third test, and it’s better they die now than after a weekend.
3. Step two: give it enough context to not be stupid
Once you know the job, the next question is what it needs to know to do that job the way you would.
Not everything about your business… just what this job needs.
For my Monday check it’s four lines. Here’s the difference between not enough and enough:
NOT ENOUGH
Check my site for new posts.
ENOUGH
Site: simplestartup.net
"Promoted" means three things: an X post, a Facebook post,
and a newsletter send
Social lives in Buffer. The newsletter is FluentCRM.
Nothing publishes without my approval.
The first version assumes the assistant knows which site, what counts, and where to look. It doesn’t. The second version fits in a breath and removes every guess.
And note what is not in there: what I sell, who I write for, my voice rules, any of it. If I pasted my whole business brief the result would be worse, not better, because the four lines that matter would be buried in four hundred that don’t.
More context is not better context. Ten relevant lines beat four thousand words of background, every time.
I’ve written separately about how to give an assistant proper context about your business and why it resets every session no matter how good your file is. The short version for this post: write it down once, store the slow-moving parts, paste the fast-moving parts, and don’t confuse having a document with having a system.
4. Step three: give it a trigger that isn’t you
This is the step that turns a chat window into an agent, and it’s where nearly everyone stops.
Most people’s “AI agent” is a conversation they open when they think of it. Which means it runs when they’re already on top of things… and goes quiet in exactly the weeks they needed it, because the weeks you forget are the weeks you were busy.
If you have to remember to use it, it isn’t an agent yet.
A trigger is anything that starts the job without your involvement. A schedule, most simply… or an event: something published, a form submitted, a payment landed, a date passed.
The distinction matters more than it sounds:
| A chat window | An agent | |
|---|---|---|
| Who starts it | You, when you remember | A schedule or an event |
| When it runs | Good weeks | Every week |
| What it knows | Whatever you paste | What you gave it, every time |
| How you know it ran | You were there | It tells you |
Look at row two. That’s the whole thing.
The tool that made this real for me was the least glamorous one available: a scheduled task. Mine fires every Monday at nine, Toronto time, and asks the same question. Setting it up took less time than deciding what it should ask.
5. Step four: decide where you stay in the loop
An agent that acts without you is not automatically better than one that asks.
For anything that reaches another human, I stay in the loop. Social posts, emails to the list, edits to live pages.
Agents draft, I approve.
For anything that only reaches me… reports, checks, drafts, summaries, research… it runs on its own and I read the result.
That’s the line, and it’s worth drawing deliberately rather than discovering it. Ask one question of each job: if this goes wrong, does anyone but me find out? If yes, you approve. If no, let it run.
Write the rule down where the agent reads it, including the cases that feel too obvious to state. An assistant that has never been told to ask will reasonably treat shipping as part of finishing.
6. The brief, before you open any tool
Everything above fits on half a page. This is what I write before touching a platform, and it takes about ten minutes.
AGENT BRIEF
Job: [what it does, in one sentence]
Runs: [when: a schedule or an event, never "when I think of it"]
Done looks like: [how you'll know it finished correctly]
Needs to know: [the 5-10 lines of context for THIS job only]
Can decide: [what it's allowed to do without asking]
Must ask: [what always comes to me first]
Tells me: [where the result lands, and whether I hear about it if nothing happened]
That last line is easy to miss and it bites. An agent that only speaks up when something interesting happens is indistinguishable from an agent that’s silently broken. Mine tells me either way, even if the answer is “nothing published last week, nothing to promote.”
Fill this in and the tool question becomes small. Almost anything runs this.
The brief is the hard part… and it’s the part nobody sells you.
7. The one I actually built, step by step
Enough theory. Here is the whole thing, exactly as it is set up, in the tool I actually use.
I build these as scheduled tasks in Claude. Same idea works in ChatGPT or anywhere else that can run a prompt on a timer… the mechanics differ, the four decisions do not.
Total setup time was about twenty minutes, and nineteen of those were spent deciding what it should say.
Step 1: The job, run against the four tests
Every Monday, check what published last week and whether it actually got promoted.
Repeats on a nameable schedule? Weekly, Mondays.
Has a definition of done? Yes. Every post from last week has a status against each channel.
Have I done it by hand? Dozens of times, badly and inconsistently… which is exactly why I can tell a good answer from a plausible one.
Near money or near the list? The newsletter check is. A post that never reaches the list is a post that only earns once, and that was the one I kept forgetting.
Step 2: Understand what the scheduled run can actually see
This is the part that trips everyone up, and it is worth slowing down for.
A scheduled task does not run inside the conversation where you created it. Every firing starts a completely fresh session. It has no memory of what you were discussing when you set it up, no idea which site you mean, no idea what you consider “promoted.”
What it does have: the prompt you saved, and whatever tools are connected to your account.
So this prompt, which would work perfectly in a conversation, is useless on a schedule:
Check if last week's post got promoted.
Last week’s post on which site? Promoted where? Compared to what? In chat it works, because the assistant can see everything you said in the last twenty minutes. On Monday morning it has none of that.
Step 3: Write the prompt for a stranger
Here is the actual text I saved. Not a summary of it… the thing itself.
Every Monday, check simplestartup.net for posts published in the
last 7 days.
For each post found, check Buffer for matching posts on X and
Facebook, and check whether a newsletter campaign went out for it.
Report back in this format:
Post title and publish date
X: published / still a draft / nothing found
Facebook: published / still a draft / nothing found
Newsletter: sent / not sent
If nothing was published last week, say so anyway. Do not skip
the message.
Read only. Do not publish, schedule or send anything.
Four things in there are doing real work, and none of them are clever.
It names the site. A fresh session cannot guess.
It defines “promoted” as three specific checks, so the answer is never a judgement call.
It says what to do when there is nothing to report. Without that line the agent goes silent on quiet weeks, and a silent agent is indistinguishable from a broken one.
It ends with a read-only instruction. This one is a habit rather than a necessity here, but it costs a line and it means the task can never surprise me.
Step 4: Set the schedule and check the connections
Monday, nine in the morning, Toronto time. That is the whole configuration.
Then the bit people skip: the run can only reach the tools connected to your account. Mine needs to see the site and it needs to see Buffer. If either connection drops, the task still fires and still writes to me, and it will cheerfully report finding nothing.
Which is worth knowing before you trust it.
Step 5: Fire it once by hand before you believe it
Do not wait until Monday to find out whether it works.
I ran mine manually the moment I saved it, and the first version came back with the wrong week… it read the current week rather than the previous one, because “last 7 days” and “last week” are not the same thing on a Monday morning. One line changed. Ran it again. Correct.
Ten minutes of testing on a Tuesday beats discovering it on the fourth Monday, by which point you have stopped reading the message because it has always been wrong.
What actually lands
That second message is the one people leave out when they build these, and it is the one that keeps the thing trustworthy.
The honest count
That is my whole fleet.
The publishing work is bigger and more useful in the moment… drafting, building the post, wiring the opt-in, checking the page rendered properly afterwards. It has context, a definition of done, and it makes real decisions.
But I start it. Every time.
Which means that by the test in this post, it is not an agent. It is very good assistance. I could put a trigger on it and I have not, because that job genuinely begins with me deciding what to write.
So the count is one.
I am telling you the number because the number is the point. You are not behind for having zero, and you will not be ahead for having twelve. One job that runs without you changes more than ten that wait for you to remember them.
8. Now write yours
Before you read the next bit, three lines. It takes two minutes and it’s the difference between finishing this post and starting something.
The job that annoys me most often:
When it happens:
How I'd know it finished correctly:
If the second line is blank, you picked a topic rather than a job. Go back to the table and take one of the six.
If the third line is blank, that’s the real work of this whole post, and it’s worth sitting with for a minute. You cannot check an answer you haven’t defined.
Everything after that is typing.
The reframe
The pitch for agents is a team of digital employees. It’s a good image and it sells software, and for one person I think it aims at the wrong thing entirely.
You don’t need staff.
You need a handful of things that keep happening while your attention is somewhere else… which is a much smaller and much more achievable idea.
That’s the real shift, and it’s quieter than the marketing suggests. Not that you can produce more. That something finally keeps happening on the weeks you drop the ball, and everyone drops the ball, and the difference between businesses that compound and businesses that stall is mostly what survives those weeks.
My Monday check cost me twenty minutes to set up and it is genuinely the dullest thing I own. It is also the only one of these I have never once thought about cancelling.
Pick the one job you’ve done by hand more times than you’d like to admit. Write the seven-line brief for it. Then give it a day and a time.
That’s your first agent. The tool you use to run it is the least interesting decision you’ll make all week.
If you want the templates and AI workflows I use to run this business solo, they’re in the Free Startup Toolkit.
FAQ
What is an AI agent, in plain terms?
Software that does a job on its own, starting from a trigger rather than from you opening a window, using context you gave it in advance. The trigger is the part that separates an agent from a chat assistant. If you have to remember to start it, you have a very capable assistant rather than an agent.
What should my first AI agent do?
Something that repeats on a nameable schedule, has a clear finish, and that you’ve done by hand enough times to spot a bad result. Boring beats impressive here, because anything you’ll actually use every week is boring by definition. If two jobs qualify, pick the one closest to a customer, a sale, or your email list.
Do I need to know how to code to build an AI agent?
No, and that’s genuinely changed. The hard part was never the building, it’s knowing the job well enough to specify it. Write the brief first: the job, the trigger, what done looks like, what it can decide alone, and what always comes to you. Most no-code platforms will run that fine.
How many AI agents should a one-person business have?
Fewer than the marketing implies. One job that runs without you changes more than ten that need you to start them. Add the second only when the first has run unattended for a month and you’ve stopped checking it.
Why do most AI agents get abandoned?
Usually because the job was chosen to suit the tool rather than the other way round, so it was never something the person actually needed doing. The second most common reason is no trigger, which means it only runs in the weeks you were already on top of things and goes quiet in the weeks you weren’t.



