A ten-minute test for AI at your job, done once in full

What came back

Migration testing completed ahead of schedule. The team is on track for the 12 Septemberwrong go-live, with the remaining budget of forty thousand dollarsinternal sufficient to cover the final phase.
Fluent, specific, and two of its facts should not have reached a client. Reading it again is not how you find that out.

The test is three steps and the first one is the step people skip. Write down what a good output looks like before you open the tool. Then run the task. Then mark what comes back against what you wrote, rather than against how it reads.

Described that way it sounds obvious. It is worth watching it happen once, because the value of the first step only shows up at the third.

The task

A weekly status report to a client, written every Friday by a project manager at a small consultancy. Half a page. What was done, what is next, anything that needs a decision. It is the kind of task an AI tool is obviously good at, which is what makes it a fair test.

Step one: the standard, written first

Three lines, written before the tool was opened. What it must include, what it must not, and the one mistake that would embarrass me in front of the client.

  • Must include: every workstream by name, even the ones with nothing to report, because silence on one reads as trouble.
  • Must not include: anything about the client's staff, and no internal cost figures.
  • The embarrassing mistake: a date or a dollar figure that is wrong. The client checks those and nothing else.

That took two minutes and it is the whole reason the rest of the test works. A standard written after you have read the output is not a standard, it is a reaction to what you were given.

Step two: run it

The week's notes went in, roughly as they were, with the client name removed. What came back was well organised, correctly structured, and pleasant to read.

Migration testing completed ahead of schedule. The team is on track for the 12 September go-live, with the remaining budget of forty thousand dollars sufficient to cover the final phase.
The tool, first attempt

Step three: mark it against the standard

Against line one it failed quietly. Two of five workstreams were missing, because the notes that week barely mentioned them, and the tool wrote about what it was given rather than what the report owed the reader.

Against line two it failed outright. The remaining budget is an internal figure and it went in without being asked for.

Against line three it failed in the way that matters. The go-live date had moved to 19 September earlier that week. The note saying so was in the source material. The sentence reads with complete confidence and the date is wrong, which is the specific failure the standard existed to catch.

Read without a standard, that paragraph is good. It is fluent, it is specific, and one of its three facts would have gone to a client wrong.

What the test is actually measuring

Not whether the tool is any good. It is measuring whether you can tell, which is a different question and the one your job now asks of you. Confidence in an output is not evidence about the output, and the only defence against grading on fluency is a standard you wrote before fluency was available to you.

Do it for five tasks and write down what you found. The job, the standard you set, what happened, and what you would do about it. That document is a findings note, and it is the first thing you produce in the Foundation of AI course, on your own work rather than a case study.

Articles on using AI at work, from Aiydo