Where scripts get stuck is usually reading, not tapping
Anyone who has automated something with a script has met this: the button is clearly there, and the script will not press it.
The cause is rarely locating. It is judgement — what the button looks like, when it should be pressed, and which path to take afterwards, all written into code in advance.
A concrete example: an order list in some back end, and you want the orders whose status shows an anomaly.
A fixed script enumerates every possible anomalous status, then writes a branch for each. That is laborious to write, and any new status means going back into the code.
A person does it differently: scan the list and pick out the ones that look wrong. No enumeration, because a person can read.
This article is about which jobs fall into that category, and how to hand them over.
Three traits
To tell whether a job is a look-then-decide job, check three things.
- It requires reading. What happens next depends on what the screen says — a number, a status, some text — not on where a button sits.
- It branches. It is not a straight line. Reading the content sends you down different paths.
- It identifies rather than locates. You are after the row that says something specific, not the point at row 320, column 65.
A job with all three is awkward to script. A job with two is worth reconsidering.
Three categories that fit
Conditional triggers
Act only when a condition is met.
Order when the price drops below a value. Screenshot when stock shows zero. Process in bulk once unread messages pass ten. The characteristic is that most of the time nothing happens.
A fixed script struggles here: either it watches constantly, which wastes resources, or it checks on a timer and misses the moment. Looking first suits it naturally: check, skip if unmet, act if met.
Content filtering
Pick qualifying items out of a list.
Comments containing a keyword, orders in a particular state, messages not yet answered. The key is that the list differs every time. You cannot write a branch per item; you can only look through and pick.
Status checks
Act only on anomalies.
An account showing abnormal, a device offline, a task failed. Similar to conditional triggers, except it requires understanding what an anomaly looks like rather than comparing a number.
A worked example
Back to picking out anomalous orders.
Described, it reads roughly like this:
Open the order list and read each row from top to bottom. If a row’s status is neither completed nor awaiting dispatch, save a screenshot of that row and note the order number. When you reach the end, put the screenshots and order numbers into one table.
There is no tap instruction in that paragraph. You said what to read, what to pick, and where it goes.
It works because it reads content, so a new status introduced by the platform needs no edit from you. It sees the unfamiliar text and picks it out anyway.
A fixed script cannot do that. It only knows the statuses you told it about, and anything else may simply be missed — often the very one that needed attention.
What to do about misreads
This is where such a capability most needs guarding.
A misread is not fatal; acting on a misread is. Reading 1,299 as 129 and then ordering at that price is.
So add a brake: stop when the value read is far from expectation.
Make the boundary explicit in the description:
If the price read differs from the last recorded value by more than half, do not act. Save a screenshot and wait for me to confirm.
One extra sentence, and it converts a misread from a potential loss into a prompt for you to look.
What does not fit
One line has to be drawn clearly: judgement without a clear rule should stay with people.
Is this comment negative, is this lead worth pursuing, is this photo good enough. Ask ten people and you may get seven answers; a machine gives you the plausible one, and you may not notice when it is wrong.
The test is simple: can you state the rule in one sentence? If yes, hand it over. If not, leave it.
How to start
Pick the shortest look, then do one thing flow.
Open a page, check whether the price changed, screenshot if it did. Three steps.
The shorter the flow, the easier it is to judge whether it reads accurately — a misread in a three-step flow is obvious, whereas in a twenty-step flow you will blame a later step.
Add branches once it runs. The order I would suggest: first “skip when the condition is not met”, then “stop on an anomaly”, and only then “actions that modify something”.
The software is completely free and runs on your own computer. For a related idea, read turning phone screenshots into a spreadsheet; for where the capability stops, see what AI phone control is.
Frequently asked questions
- What counts as a look-then-decide job?
- The screen shows some content, you have to read it before you know where to tap next. Order when the price drops below a value, screenshot when a status looks wrong, continue when a list has new entries. The difficulty is reading, not locating.
- How is that different from an ordinary script?
- A script follows a fixed sequence: tap here, wait three seconds, tap there. It does not read, so every branch has to be written in advance. Looking first means deciding at that moment based on what is on screen.
- What makes it possible?
- Reading the current screen, either structured page text or text inside an image, and then judging. That is also why it survives interface changes: the basis for the decision is content, not position.
- Which jobs fit best?
- Three: conditional triggers where something only happens past a threshold, content filtering where you pick qualifying items out of a list, and status checks where you act only on anomalies. All three are common and all three are awkward to script.
- Which jobs do not fit?
- Anything needing subjective judgement: whether a comment counts as negative, whether a lead is worth pursuing. There is no clear rule, and a machine will produce a plausible wrong answer.
- What if it reads something wrong?
- A misread is not the danger; acting on it is. So add a brake: when the value read is wildly different from expectation, stop rather than pressing on.
- Is it slower than a fixed script?
- Slightly, because of the extra read and judge. What that buys is not rewriting when the interface changes, and being able to handle branches. In most business work that trade is worth it.
- How do I start?
- Pick the shortest look-then-do-one-thing flow. Open a page, check whether the price changed, screenshot if it did. The shorter the flow, the easier it is to tell whether it reads accurately.
Got work like this?
Let it look first, then decide
Software is completely free and runs on your own computer. Start with one small job.