How the old work quiz works
This page explains the old 86-question quiz, its scores, and how it chooses which jobs to try first.
The interactive version administers this instrument and generates the output. This document specifies it: the question bank with the purpose of every item, the scoring constants, the phase assignment rules, the eight divergence checks, the thirteen prerequisite triggers, and the structure of the brief. Read it if you need to defend the output to somebody, adapt the instrument to a domain it does not currently handle, or satisfy yourself that the ordering is derived rather than asserted.
The one-paragraph version
The respondent lists the jobs they would hand over and scores each on six dimensions. Three of those dimensions (verifiability, reversibility and reach) decide which of five phases a job belongs in; the other three (frequency, duration and readiness) decide its rank inside that phase. Preference is collected but never enters the formula. The instrument then compares what the respondent said they wanted against what the scoring produced, and reports every disagreement. The output is a briefing document written to be read by a Bot at install, not by a human.
Purpose
What the instrument is trying to produce, and the two things it deliberately will not do.
1.1What it is for
Grok Bot sets itself up badly from a description of what somebody wants and well from a description of how they work. The gap between those two is where almost all first-week failure lives. This instrument exists to produce the second thing.
Its output is a single document a person pastes into their first Bot before creating any others. That document carries four things a Bot cannot infer and will otherwise invent: the operator's context and vocabulary, the boundaries in the operator's own words, the order in which capability should be handed over, and an instruction set telling the Bot to propose rather than act.
The order is the part that justifies the instrument. Everything else could be captured in a conversation. Sequencing cannot, because the correct sequence depends on properties of each job that people systematically misjudge about their own work.
1.2What it refuses to do
It will not rank by importance. The respondent's stated priority is collected in two places and used only to detect disagreement. It never influences the build order. A job that would transform the business but cannot be verified quickly still sorts below a trivial job that can, because the purpose of the first build is not value. It is evidence.
It will not raise the ceiling on request. Every generated brief starts at draft-only for four weeks regardless of the autonomy answer, and every irreversible action stays gated permanently. When a respondent asks for sending or transacting, the instrument records the request, fires a divergence explaining the refusal, and overrides it in the brief. This is a deliberate design position rather than an oversight: approvals in this product gate proposed actions and do not reverse completed ones, and this earlier assessment used a fixed review policy. See the current security records for available Enterprise audit controls.
The thesis
One claim, and the three properties it rests on.
2.1Why order beats preference
People abandon agent tools in week two, and almost never because the tool was incapable. They abandon them because they handed over something important on day two, it went eighty percent right unsupervised, and they had no way to tell whether that was competence or luck. Having no way to tell is the actual failure. The work was never the problem.
That failure is entirely predictable from properties of the task chosen, and those properties are knowable in advance. A job you can check in thirty seconds produces evidence every time it runs. A job that takes an hour to verify produces evidence roughly never, because nobody spends the hour. Trust accumulates from the first kind and cannot accumulate from the second, whatever the second is worth.
So the instrument sorts by evidence yield rather than by value, and tells the respondent plainly that it is doing so. The most valuable job in the set frequently sorts fourth. That is not an error to apologise for; it is the finding.
2.2The three dimensions that decide a phase
Verifiability
How fast the operator can tell the output is wrong. Determines whether a job can produce evidence at all. A job nobody can check quickly cannot be scheduled, however safe it is.
Reversibility
What happens when it is wrong. The line is not drawn on importance; it is drawn on whether a mistake can be taken back in under a minute. Anything reaching a customer or touching money is not reversible, whatever its size.
Reach
How far into connected systems the job goes. Files are free. Reading an internal system is nearly free. Writing to an external system is where blast radius, datacentre blocking and cost all appear at once.
Three further dimensions (frequency, duration and whether an example of good output exists) decide rank within a phase but never move a job between phases. This separation is the core of the design: value competes only against comparable risk.
The scoring model
Every constant, the formula, the assignment rules and a worked example. Nothing here is hidden from the respondent.
3.1Constants
Six dimensions are collected per job. Four of them carry numeric weights:
| Dimension | Answer | Value | Used for |
|---|---|---|---|
| Frequency | every day | 22 | Runs per month, converts minutes into hours a year |
| weekly | 4.33 | ||
| monthly | 1 | ||
| quarterly | 0.33 | ||
| ad hoc | 0.5 | ||
| Verifiability | checkable in 30 seconds | 3 | Phase gate and Phase 1 eligibility |
| checkable in five minutes | 2 | ||
| takes an hour to verify | 1 | ||
| hard to verify at all | 0 | ||
| Reversibility | no consequence if wrong | 3 | Phase gate; a zero forces Phase 4 |
| I just redo it | 3 | ||
| someone internal sees it | 2 | ||
| a customer sees it | 0 | ||
| money or legal exposure | 0 | ||
| Reach | files only | 3 | Phase gate; a zero forces Phase 4, a one forces Phase 3 |
| reads an internal system | 3 | ||
| writes to an internal system | 2 | ||
| reads an external system | 2 | ||
| writes to an external system | 1 | ||
| moves money | 0 |
The two remaining dimensions are duration (minutes per run, entered as a number and floored at one) and readiness (whether an example of the job done well exists, worth one or zero).
3.2The formula
value = annual hours × (0.8 + 0.2 × readiness)
Rank within a phase is by value, descending.
Rank never crosses a phase boundary.
The readiness multiplier is deliberately small. Its job is to break ties in favour of work the operator can already demonstrate, not to reorder the list. A job worth two hundred hours a year with no example still outranks a job worth forty with one.
Frequency is expressed as runs per month, so an ad-hoc job is scored at 0.5 rather than zero. Ad-hoc work is real work; it simply does not compound, and the low weight reflects that a job which happens unpredictably cannot anchor a routine.
3.3Phase assignment
Evaluated in order. The first matching rule wins.
| Rule | Condition | Phase | Why |
|---|---|---|---|
| 1 | reversibility = 0 or reach = 0 | 4 | A customer sees it, or money moves. Approvals gate proposals and do not reverse completed work, so these stay behind a human indefinitely rather than for a probation period. |
| 2 | verifiability = 0 | 3 | Cannot produce evidence. Scheduling this creates a routine whose failure is indistinguishable from its success. |
| 3 | reach ≤ 1 | 3 | Writes to an external system. Recoverable in principle, but the blast radius and the cost both justify waiting until the operator has calibration. |
| 4 | verifiability ≥ 2 and reversibility ≥ 2 and reach ≥ 2 | 1 | Checkable in five minutes or less, harmless if wrong, and reaching no further than an internal read. This is the only shape that earns trust quickly. |
| 5 | everything else | 2 | Safe, but either slower to verify or reaching further than a first build should. |
Phase 0 is not produced by this rule set. It is assembled separately from the prerequisite triggers in section 5.2 and contains no jobs at all, only conditions that must be true before any Bot is created.
The asymmetry is intentional
Rules 1 to 3 can only push a job later. Rule 4 is the only rule that pulls one forward, and it requires all three safety dimensions simultaneously. An instrument that let a high-value job argue its way into Phase 1 would reproduce exactly the failure it exists to prevent.
3.4A worked example
Six jobs from a two-person trade business, scored as the respondent entered them.
| Job | Freq | Mins | V | R | K | Hours/yr | Phase |
|---|---|---|---|---|---|---|---|
| Morning inbox triage | daily | 25 | 3 | 3 | 3 | 110 | 1 |
| Job board tidy | daily | 12 | 3 | 3 | 2 | 53 | 1 |
| Supplier price comparison | weekly | 45 | 2 | 3 | 2 | 39 | 1 |
| Chasing supplier invoices | weekly | 40 | 1 | 2 | 1 | 35 | 3 |
| Quote drafting | daily | 35 | 2 | 0 | 2 | 154 | 4 |
| Certificate renewals sweep | monthly | 90 | 0 | 0 | 1 | 18 | 4 |
Quote drafting is the largest single cost in the set at 154 hours a year, roughly forty percent more than anything else. It sorts last, because a customer sees the output. Morning inbox triage, worth less, sorts first because it is daily, checkable at a glance and harmless when wrong.
The respondent listed chasing supplier invoices first, which is a reliable signal that it is the most irritating item rather than the most valuable one. It scores into Phase 3. The instrument reports all three of these observations as disagreements rather than quietly reordering, because the respondent needs to understand the reasoning to accept the sequence.
The instrument
All 86 questions, what each one is for, and what it changes downstream. Ten of them are behavioural and are treated separately in section 4.2.
4.1The question bank
Sections 00 to 08 are asked of everyone. Sections M1 to M7 open conditionally, and the gates are listed in section 5.3. Seventy-two of the 86 items are required; the optional ones are those where a blank answer is itself informative.
00 · Before anything else 5 questions
Gate. Any of these can stop the exercise or add a Phase 0 prerequisite. Asked first so nobody spends half an hour planning something they cannot run.
| Question | Type | What it does |
|---|---|---|
| Which of these do you have?Grok Bot rides on a plan you already hold. There is no standalone subscription.A paid Cursor plan (Pro, Pro+, Ultra or Teams) · An individual SuperGrok, SuperGrok Plus or Heavy plan · X Premium+ · None of these yet · Not sure | Single choicerequired | Eligibility. A "none" or "unsure" answer adds a Phase 0 item rather than failing the respondent out. |
| Is your Cursor account on Legacy Privacy Mode?This one is a hard block, not a degraded experience. Grok Bot requires cloud data storage and will refuse to start.No, or my account is new · Yes · I do not know | Single choicerequired | Hard block. Legacy Privacy Mode prevents Grok Bot starting at all. An unsure answer is treated as a prerequisite rather than a pass. |
| What will you run the desktop app on?There is no web version. grok.com is not a Grok Bot client.macOS · Windows · Linux · iPhone as well · None of these | Multiple choicerequired | Context for the brief. |
| Does anyone else have to approve tools that hold your logins?An employer, a partner, a compliance function, a client contract, a professional body.No, it is my decision · Yes, and I have approval · Yes, and I do not have it yet · I am not sure | Single choicerequired | The question people skip and then regret. "Yes and I do not have it" is a hard Phase 0 item. |
| Will this work touch regulated, client-confidential or personally identifying data?Health records, legal matters, financial records for others, candidate data, children's data, anything under a confidentiality obligation.No · Possibly, at the edges · Yes, centrally | Single choicerequired | Module gate for M1, and one of two questions that can push the entire exercise toward a qualified review. |
01 · You 10 questions
Operator assessment. Predicts review behaviour and supplies the raw material for voice. Four of these are behavioural rather than opinion questions.
| Question | Type | What it does |
|---|---|---|
| What should the Bots call you?First name is fine. It goes into the brief. | Short textrequired | Context for the brief. |
| What is your role, in the words you would use to another person in your industry?Not your job title on paper. What you actually do all day. | Short textrequired | Context for the brief. |
| Your time zone and normal working hoursEvery routine will be scheduled against this. Getting it wrong is the most common cause of overnight bills for work nobody reads. | Short textrequired | Context for the brief. |
| How comfortable are you with technical setup?This changes how much the brief explains versus assumes.I use apps, I do not configure them · I can follow instructions and fix things when they break · I write scripts or code · I build software professionally | Single choicerequired | Context for the brief. |
| When you last delegated something important to a person, what happened?This is the best available predictor of how you will behave with a Bot, and it is asked for that reason.I have never really delegated anything important · I took it back and did it myself · I checked constantly and it worked out · I left them to it and it was mostly fine · I set up a way to check the output, not the process | Single choicerequired · behavioural | The strongest single predictor in the instrument. "I took it back" combined with any autonomy above draft-only fires a divergence. |
| What did you tell yourself you would automate or systematise three months ago, and then did not?Be specific, and include why it stalled. This is often the single most useful answer in the whole questionnaire. | Long textrequired · behavioural | Behavioural. Past inaction predicts future inaction far better than stated intent, and the stated reason usually names the real obstacle. |
| Which task takes you longer than you would happily admit?Nobody sees this but you. It is usually the highest-value thing in the list and it is almost never the thing people put first. | Long textbehavioural | Behavioural. Shame suppresses reporting, so this is asked explicitly and privately. It routinely surfaces the highest-value task in the set. |
| If you were unreachable for two weeks starting tomorrow, what breaks first?Name the specific thing and who would notice. | Long textrequired · behavioural | Behavioural. Reveals single points of failure the respondent has normalised. |
| Realistically, how often will you review what a Bot did?Answer for the version of you that exists in week six, not week one.Every day, I will read everything · A few times a week · Once a week if it is on the calendar · Honestly, only when something looks wrong | Single choicerequired | Calibration. Asked about week six rather than week one, because week one is not the problem. "Rarely" plus wide autonomy fires a divergence. |
| When something you produce goes out wrong, what usually happens?This calibrates how tight the boundaries need to be.I fix it, nobody notices · Mild awkwardness, an apology · It costs money or a client · Regulatory, legal or safety consequences | Single choicerequired | Sets how tight the boundary language needs to be. |
02 · Your actual week 8 questions
Evidence collection. Surfaces work the respondent would not think to list, and produces the reported-hours figure used in the under-listing check.
| Question | Type | What it does |
|---|---|---|
| How many hours a week do you spend on work that produces nothing new?Copying between systems, chasing, formatting, re-reading, logging in somewhere to check one number.Under 5 · 5 to 10 · 10 to 20 · More than 20 | Single choicerequired | Produces the claimed-waste figure. Compared against the sum of listed task hours to detect under-listing. |
| What is the first thing you do at your desk, before real work starts?Describe the actual sequence, including the tabs. | Long textrequired | Context for the brief. |
| What interrupts you most often, and where does it come from?Channel and typical content. | Long textrequired | Context for the brief. |
| What is the last thing you did that you had already done before, in exactly the same way?Think about yesterday and the day before. | Long textrequired · behavioural | Behavioural. Asks for a specific recent instance rather than a general pattern, which defeats the tendency to answer aspirationally. |
| What do you re-do because somebody else did it wrong the first time?Leave blank if you work alone and nothing applies. | Long textbehavioural | Behavioural. Rework caused by others is invisible in most self-assessment and is often the easiest thing to remove. |
| What are you usually waiting on, and who has it?Bottlenecks outside your control are candidates for a chasing routine. | Long text | Context for the brief. |
| What do you intend to check regularly, but do not?Prices, reviews, renewals, arrears, deadlines, competitor moves, your own numbers. | Long textrequired · behavioural | Behavioural. The gap between intended and actual monitoring is where watcher Bots earn their place. |
| Is there a date or event driving this?A launch, a season, a hire you are trying not to make, an audit.No, this is exploratory · Yes, within a month · Yes, within a quarter · A general sense that this year is the year | Single choice | Context for the brief. |
03 · The business 8 questions
Context and routing. Sets the vocabulary the brief writes in, seeds the task suggestions, and captures the stated priority used in the divergence check.
| Question | Type | What it does |
|---|---|---|
| Business nameUsed throughout the brief so the charters are paste-ready. | Short textrequired | Context for the brief. |
| What does the business actually do, and for whom?Two or three sentences, in your own words. This becomes the Bot's shared knowledge and shapes everything it writes. | Long textrequired | Context for the brief. |
| Which of these is closest?Used to seed the task suggestions and pick the right vocabulary.Solo professional services or consulting · Agency or studio with clients · Software or SaaS · Ecommerce or physical product · Local, field or trade services · Content, media or education · Regulated professional practice (legal, accounting, medical, financial) · Property, lettings or real estate · Nonprofit or public sector · I am an employee improving my own work · Something else | Single choicerequired | Context for the brief. |
| How many people?Determines the roster size the brief recommends.Just me · 2 to 10 · 11 to 50 · More than 50 | Single choicerequired | Context for the brief. |
| How is the customer base shaped?Separation requirements depend on this.A handful of named clients or accounts · Many clients, each with their own systems or portals · Lots of small customers, no per-customer systems · No external customers | Single choicerequired | Context for the brief. |
| Which of these would move the business most in the next 90 days?Your judgement, not a formula. This is one half of the divergence check at the end.More qualified leads · Closing more of what we already have · Delivering faster or with fewer mistakes · Keeping customers we already have · Spending less to do the same work · Getting my own time back | Single choicerequired | Stated priority. One half of the divergence check; the scored task order is the other half. |
| Is the work seasonal or spiky?Affects whether routines should be paused between peaks.Fairly flat · Predictable weekly rhythm · Strong seasonal peaks · Unpredictable | Single choice | Context for the brief. |
| Who do you actually lose to, and where would you check on them?Leave blank if this is not relevant. Names and URLs are fine. | Short text | Context for the brief. |
04 · Your systems 6 questions
Feasibility and cost. Determines where each job sits on the connection ladder, which is the single largest driver of both reliability and spend.
| Question | Type | What it does |
|---|---|---|
| List the systems you personally open in a normal weekOne per line. Include the ugly internal ones and the supplier portals, because those are where the value is. | Long textrequired | Context for the brief. |
| Which of those have a proper integration or connector available?If you are not sure, leave it blank and the brief will tell the Bot to find out first. | Long text | Context for the brief. |
| Which are websites you log into by hand, with no integration?These are the highest-value targets and the most fragile. Name them specifically. | Long textrequired | Highest-value and most fragile targets. Drives both the connection ladder and the datacentre-blocking warning. |
| Have you ever seen any of these block automated or unusual traffic?Bots reach the web through shared datacentre addresses, and some sites refuse them after a perfectly correct login.Not that I know of · Yes, at least one · No idea | Single choice | Context for the brief. |
| Can you create limited service accounts in your main systems?A separate login scoped to the minimum the work needs. This is the only real security control available, because Grok Bot cannot give one Bot less reach than another.Yes, in most of them · In some · No, there is one login and it is mine · I do not know | Single choicerequired | The only real security control available, because the product cannot scope one Bot below another. Anything but "yes" adds a Phase 0 item. |
| Where do working files live?The Bot needs one durable place to keep project state.A cloud drive · My own machine · Both, inconsistently · Scattered across email and apps | Single choicerequired | Determines whether handoffs can travel as file paths. "Nowhere consistent" adds a Phase 0 item. |
05 · The work itself 3 questions
The engine. Everything else contextualises; this section produces the ordering.
| Question | Type | What it does |
|---|---|---|
| The jobs you would hand overAdd each one, then score it. Score honestly rather than optimistically: a task you mark as easy to verify will be scheduled early, and if that was wishful thinking you will find out the expensive way. | Repeating scored rowsrequired | The scoring engine. Six dimensions per job, three of which decide the phase and three of which decide the rank inside it. |
| Which one of those did you write down first?Type its name. First-listed is usually what is most annoying rather than what is most valuable, and the brief will say so if the scoring disagrees. | Short textrequired · behavioural | Behavioural. First-listed is a proxy for most-annoying. Compared against the top-ranked Phase 1 job to fire the primary divergence. |
| Which one would you least want to check the output of?Type its name, or leave blank. A task you would not want to verify is a task you should not automate yet, whatever it scores. | Short textbehavioural | Behavioural. A job the respondent would not want to verify is a job that will fail silently, whatever it scored. |
06 · Voice and standards 5 questions
Output quality. Decides whether drafts read like the respondent or like a model.
| Question | Type | What it does |
|---|---|---|
| What will the Bot write on your behalf?Everything ticked here needs a voice profile before it drafts anything.Email · Internal notes and updates · Reports and summaries · Marketing or social content · Proposals or quotes · Customer support replies · Documentation or process · Nothing customer-facing | Multiple choicerequired | Context for the brief. |
| Do you have a body of your own writing it could learn from?Sent mail counts. So does anything you have published.Yes, years of sent mail · Some, if I go and find it · Yes, published work · Not really | Single choicerequired | Decides whether voice extraction is possible at install or has to be a prerequisite. |
| What should it never do in your writing?Your tells. Words, punctuation, openings, moves that are not you. | Long text | Context for the brief. |
| Can you point at three examples of your own work that are genuinely good?Demonstrated standards beat described ones by a wide margin.Yes, I know exactly where they are · I could find them · No | Single choicerequired | Demonstrated standards beat described ones. Anything but "yes" adds a Phase 0 item when the Bot will write. |
| For your most common deliverable, what makes it acceptable?Write checks, not adjectives. "Every figure has a source and a date" rather than "professional". | Long textrequired | Context for the brief. |
07 · Boundaries 6 questions
Boundary capture, in their own words rather than a template. Produces the Red list in the brief.
| Question | Type | What it does |
|---|---|---|
| What must never happen without you?Write it as a list. Be specific to your business rather than generic. | Long textrequired | Captured in the respondent own words rather than selected from a list, because a boundary someone wrote is a boundary they remember. |
| Will any of this be near money?Payments, invoices, refunds, purchasing, payroll, brokerage.No · It will read financial data but never move anything · It will prepare payments for a human to approve · I want it to actually transact | Single choicerequired | "I want it to transact" fires a divergence. The brief overrides it to prepare-and-approve regardless. |
| Should it ever contact a person outside the business by itself?The honest answer for a beta product is almost always no.Never, everything is a draft · Only after I approve each message · Within rules I define, for low-risk messages · Yes, I want it sending | Single choicerequired | "Yes, I want it sending" fires a divergence and is deliberately overridden in the generated brief. |
| Should it be able to run commands on your own computer?This is separate from its cloud machine. The default is ask every time, and off is usually correct.No · Only with approval each time · Yes, it needs my local files | Single choicerequired | Context for the brief. |
| What is the worst thing that could plausibly happen if this goes wrong unattended overnight?Write the actual scenario, not the abstract risk. This becomes the sentence that justifies every red line in your brief. | Long textrequired · behavioural | Becomes the sentence in the brief that justifies every red line. Written by the respondent, so it survives being reread later. |
| What would make you turn the whole thing off?Decide now, while you are calm. Deciding this while annoyed produces worse decisions. | Long text | Decided while calm. The instrument asks for it now precisely because it cannot be decided well later. |
08 · Constraints and appetite 8 questions
Constraints and module gates. Six of these eight questions open or suppress a module.
| Question | Type | What it does |
|---|---|---|
| What worries you about running this?Each one opens a module later in the brief with specific countermeasures.Burning through usage and cost · Access and credentials · It will produce generic or wrong work · I will not have time to supervise it · I will not be able to tell whether it is working · Nothing much yet | Multiple choicerequired | Module gate for M7 and the driver of the cost-discipline block in the brief. |
| How much time can you give setup in the first two weeks?Be realistic. This decides how many Bots the brief tells you to create.About an hour total · Half a day · A day or so · A few hours a week, ongoing | Single choicerequired | Context for the brief. |
| What happens if usage runs past the included allowance?Review your shared account on-demand setting and spend limits.It cannot. There is a hard limit on what I will spend · Some overage is fine if it is producing · Cost is not the constraint · I do not know what the allowance is | Single choicerequired | "Hard limit" combined with three or more daily jobs fires a divergence. |
| Where do you want the leash after four weeks?Everyone starts at draft-only. This is about the destination.Still drafting. I send everything · Acting on anything reversible, asking for the rest · Wide internal autonomy, irreversible things gated | Single choicerequired | Destination, not starting point. Every generated brief starts at draft-only regardless of this answer. |
| Do you already run Bots, agents or automations?Includes other agent tools, scripted automations and scheduled scripts.No, starting fresh · A few automations, no agents · Yes, agents in another tool · Yes, Grok Bots already | Single choicerequired | Context for the brief. |
| Does software or code work fall inside what you want handled?Building, fixing, testing, reviewing or shipping software.No · Light scripting and data work · Yes, real software work | Single choicerequired | Context for the brief. |
| Does any of this touch the physical world?Stock, deliveries, vehicles, premises, appointments with a person at a place.No, it is all digital · Some scheduling or stock · Yes, centrally | Single choicerequired | Context for the brief. |
| In four weeks, what will have to be true for this to have been worth it?One sentence. A number if you can. This becomes the metric in your brief. | Long textrequired | Becomes the metric in the brief. One sentence, ideally with a number. |
M1 · Regulated or confidential data 5 questions
Opens when regulated data is in scope. Changes the data boundary in the brief and can add a Phase 0 item requiring qualified review.
| Question | Type | What it does |
|---|---|---|
| What kind of sensitive data is involved?Health or medical · Legal or matter-privileged · Financial records belonging to others · Candidate or employee records · Data about children · General personal data about customers · Client confidential or trade secret | Multiple choicerequired | Context for the brief. |
| What rule, contract or obligation governs it?Name the regime or the clause if you know it. If you do not, say so plainly. | Long textrequired | Context for the brief. |
| Could the work be done from aggregates or de-identified extracts instead?Most reporting can. This is usually the difference between a defensible setup and an indefensible one.Yes, mostly · Some of it · No, it needs the underlying records | Single choicerequired | "No, it needs underlying records" produces the strongest warning the brief contains. |
| Do you need to be able to reconstruct what happened on a given date?Review current Enterprise audit logs, Action Recording and export options against your evidence requirements.No · It would be useful · Yes, it is a requirement | Single choicerequired | "It is a requirement" adds a Phase 0 item, to define the evidence and retention required for this workflow. |
| Has anyone qualified reviewed this decision?No · Only me · Someone internal with the relevant responsibility · External counsel or a compliance professional | Single choicerequired | Context for the brief. |
M2 · Multiple clients or accounts 4 questions
Opens for many-client or agency shapes. Addresses the fact that Grok Bot cannot enforce separation between clients.
| Question | Type | What it does |
|---|---|---|
| How many clients or accounts have their own systems?2 to 5 · 6 to 20 · More than 20 | Single choicerequired | Context for the brief. |
| Do you have a contractual obligation to keep client data separated?Every Bot on one account shares one machine, one filesystem and one set of logins. Separation between Bots does not exist.No · Implied but not written · Yes, explicitly | Single choicerequired | An explicit separation obligation plus no funded seats produces a Phase 0 item and a procedural-separation warning. |
| Could you run a separate paid seat per client group?This is the only real isolation available.Yes · For the biggest ones · No, not economic | Single choicerequired | Context for the brief. |
| Which client systems would it need to log into? | Short text | Context for the brief. |
M3 · Other people 4 questions
Opens above one person. Determines whether the brief must survive being handed to somebody else.
| Question | Type | What it does |
|---|---|---|
| Who else will use, see or be affected by this?Roles rather than names. | Long textrequired | Context for the brief. |
| Are you on a Teams or Enterprise plan with an admin?Team controls change the picture considerably: network policy, connector allowlists, team rules and action recording.No, individual plan · Teams plan · Enterprise · Not sure | Single choicerequired | Context for the brief. |
| Will more than one person give instructions to the same Bot?No, just me · Two or three of us · Anyone on the team | Single choicerequired | Context for the brief. |
| If you left tomorrow, could someone else run this setup?If the answer is no, the brief needs to be a document rather than a conversation.No · With effort · Yes, it is written down | Single choicerequired | Context for the brief. |
M4 · What you already run 3 questions
Opens when something already runs. Prevents the plan duplicating or contradicting existing automation.
| Question | Type | What it does |
|---|---|---|
| What do you already run, and what does each one own? | Long textrequired | Context for the brief. |
| Which of them actually works, and which have you quietly stopped trusting?Be honest. Duplicating a broken automation in a new tool produces two broken automations. | Long textrequired | Context for the brief. |
| Where does state live in the existing setup?In conversations · In files · In a project tool or database · Nowhere durable | Single choicerequired | Context for the brief. |
M5 · Software and code 4 questions
Opens when software work is in scope. Sets the outer-loop configuration and checks for a staging environment.
| Question | Type | What it does |
|---|---|---|
| Can you issue a scoped token limited to one project?Yes · No, tokens are broad · Not sure | Single choicerequired | Context for the brief. |
| Is there a staging environment with non-production data?There is no dry run in Grok Bot. The first execution is real.Yes · Sort of · No, it is production or nothing | Single choicerequired | "Production or nothing" restricts the code Bot to reading, because there is no dry run. |
| Do you already pay for a coding agent?The strongest configuration is Grok Bot as the outer loop staging work for a coding agent that does the building, because the token cost lands on a different meter.Yes · No · I could | Single choicerequired | Context for the brief. |
| What should it be allowed to do with code?Read and explain · Reproduce reported bugs in staging · Open draft pull requests · Write and run tests · Never merge, deploy or release | Multiple choicerequired | Context for the brief. |
M6 · Physical operations 3 questions
Opens when the work touches the physical world, where failures cost journeys and appointments rather than tokens.
| Question | Type | What it does |
|---|---|---|
| What does it touch?Stock or inventory · Orders and deliveries · Scheduling people to places · Vehicles, premises or equipment · Certificates, inspections or licences | Multiple choicerequired | Context for the brief. |
| What does a wrong answer cost in the real world?A wasted journey, a missed inspection, a customer waiting in. | Long textrequired | Context for the brief. |
| Which supplier or booking portals would it use? | Short text | Context for the brief. |
M7 · Cost control 4 questions
Opens when cost is a stated concern. Ten more questions on the product mechanic that surprises people most.
| Question | Type | What it does |
|---|---|---|
| Do you know what your weekly allowance actually is?It is not published in any unit. The plan screen inside the app is the only place to see where you stand.Yes, I have looked · No | Single choicerequired | Context for the brief. |
| Is there anything that genuinely has to run overnight?Almost nothing does. Scoping routines to working hours is the single cheapest saving available.No · One or two things · Yes, monitoring that cannot wait | Single choicerequired | Context for the brief. |
| Were you planning to point it at a large history on day one?A full inbox, an archive, years of records.No · Yes, that was the plan | Single choicerequired | "Yes, that was the plan" writes an explicit instruction into the brief telling the Bot not to let them. |
| Is there anything you want it to do in a loop, over many items?Anything that loops should become a script it runs once rather than a conversation it iterates through. | Long text | Context for the brief. |
4.2The behavioural questions
Ten items do not ask what the respondent thinks. They ask what the respondent did, because stated intent is a poor predictor of behaviour and people are least accurate about exactly the things this instrument needs to know. Each is written to defeat a specific reporting bias.
| Question | Bias it defeats | What it feeds |
|---|---|---|
| When you last delegated something important to a person, what happened? | Optimism about future self. Everyone believes they will delegate well; almost nobody has evidence. | Divergence 4. "I took it back" plus any autonomy above draft-only produces the strongest warning in the output. |
| What did you tell yourself you would automate three months ago, and then did not? | Recency and intent substitution. Asking what someone wants produces a wish; asking what they failed to do produces a constraint. | Context in the brief, and frequently the true first task. The stated reason usually names the real obstacle. |
| Which task takes you longer than you would happily admit? | Shame suppression. The highest-value item is often unreported because reporting it is embarrassing. | Prompts task listing. Routinely surfaces the largest single line in the set. |
| If you were unreachable for two weeks, what breaks first? | Normalisation of single points of failure. People stop seeing dependencies they have carried for years. | Written into the brief so the Bot knows what is load-bearing. |
| What is the last thing you did that you had already done before, in exactly the same way? | Generalisation. A question about patterns gets an idealised answer; a question about yesterday gets a real one. | Task discovery. |
| What do you re-do because somebody else did it wrong? | Invisibility of rework. This work is rarely counted as work at all. | Task discovery, and a signal that the constraint may be a process rather than a Bot. |
| What do you intend to check regularly, but do not? | Intent-behaviour gap, asked directly rather than inferred. | The clearest source of watcher-Bot candidates in the instrument. |
| Which job did you write down first? | Salience. First-listed tracks irritation, not value. | Divergence 1, the primary check, comparing it against the top-ranked Phase 1 job. |
| Which job would you least want to check the output of? | Avoidance. People will not admit they intend to skip verification, but they will name the task. | Divergence 2. Any such job scoring into Phase 1 or 2 is flagged regardless of its score. |
| What is the worst thing that could plausibly happen overnight? | Abstraction. "Risk" is dismissable; a specific scenario is not. | Becomes the sentence in the brief that justifies every red line, in the respondent's own words. |
Two design rules govern these items. They ask for a specific past instance rather than a general pattern, and they are worded so that the uncomfortable answer is the easy one to give. Both are why they are placed early, before the respondent has worked out what the instrument rewards.
Derived logic
Everything the instrument concludes rather than collects: eight disagreement checks, thirteen prerequisite triggers, seven module gates.
5.1Divergence rules
These compare stated preference against scored evidence and report every conflict. They are the reason the output is persuasive rather than merely correct: a sequence somebody does not understand is a sequence they abandon.
| # | Fires when | What it says |
|---|---|---|
| 1 | The job listed first is not the top-ranked Phase 1 job | Names both, explains the winning job's frequency, verifiability and reversibility, and states which rule pushed the respondent's choice into a later phase. |
| 2 | The job the respondent would least want to verify scored into Phase 1 or 2 | A job nobody will check fails silently. Either design a check they would actually perform, or gate it regardless of score. |
| 3 | The highest-value job overall lands in Phase 3 or 4 | States the hours, states the phase, and reframes the first fortnight as making that job verifiable so it can move up. |
| 4 | Review habit is "rarely" and autonomy is "wide" | Names the combination as the failure everyone hits and asks them to change one of the two answers. |
| 5 | Last delegation ended in taking it back, and autonomy is above draft-only | Cites it as the best available predictor and argues for a longer draft-only period. |
| 6 | Listed task hours fall below 40 percent of self-reported wasted hours | Tells them the gap is where the value is hiding and sends them back to add what they omitted. |
| 7 | Hard spend limit plus three or more daily jobs | Explains step-based metering and recommends one daily routine measured for a full week before adding a second. |
| 8 | Respondent asked to transact, or to send externally | Refuses in the brief and explains why: approvals do not reverse, the session persists across every Bot, and this earlier instrument uses a fixed approval policy. |
Rules 6 and 8 are the two that most often change what somebody does. Rule 6 because under-listing is near universal and invisible without the arithmetic; rule 8 because it is the only place the instrument openly overrides a direct request.
5.2Prerequisite rules
Thirteen conditions produce Phase 0 items. None involves a Bot. Any of them left unresolved either blocks the product or removes a control the rest of the plan assumes.
| Trigger | Produces |
|---|---|
| Legacy Privacy Mode is on, or unknown | Turn it off, or check. It blocks the product entirely rather than degrading it. |
| No eligible plan, or unsure | Get onto one, or confirm on the plan screen, which is the only reliable answer per account. |
| Approval required and not held | Get it. This tool holds live sessions for every connected system. |
| No supported machine | Desktop app on macOS, Windows or Linux. There is no web version. |
| Cannot create scoped service accounts | Create them. It is the only real control, because the product cannot scope one Bot below another. |
| No durable place for files | Pick one. Handoffs that travel as file paths are cheaper and more reliable than pasted content. |
| Bot will write, and no voice corpus exists | Assemble twenty to thirty real replies. Without it every draft reads like a model. |
| Bot will write, and no examples of good work identified | Find three. Demonstrated standards beat described ones. |
| Audit reconstruction is a requirement | Build an external action log. The product has none. |
| Explicit client separation obligation, unfunded seats | Resolve it. Either fund a seat per client group or accept that separation is procedural. |
| Software in scope, no staging environment | Get one. There is no dry run; the first execution is real. |
| Regulated data central, reviewed by nobody qualified | Get a second opinion before this becomes one person's decision. |
| Regulated data needs underlying records | Strongest warning in the output. Most of that work probably should not be here. |
5.3Module gates
| Module | Opens when |
|---|---|
| M1 Regulated or confidential data | Sensitive data is in scope, even partly |
| M2 Multiple clients | Many clients with their own systems, or an agency shape |
| M3 Other people | Headcount above one |
| M4 Existing automation | Anything already runs |
| M5 Software and code | Code work is in scope at any level |
| M6 Physical operations | The work touches stock, vehicles, premises or appointments |
| M7 Cost control | Usage cost named as a concern |
Gates are re-evaluated on every relevant answer, so modules appear and disappear live rather than at a submit step. Progress is calculated against active sections only, which keeps the completion figure honest for a respondent who opens four modules.
The output
What the brief contains, in order, and why each part is where it is.
6.1Anatomy of the brief
The generated document is written to be read by a Bot rather than by a person. That changes the register: it is imperative, it front-loads constraints, and it repeats the things a model is most likely to drift from.
| Section | Why it is there, and why in that position |
|---|---|
| Instructions for you, the first Bot | First, because everything after it is context and a model that starts acting on context is the failure. Five numbered steps ending in "create one Bot, and stop." |
| Who I am | Includes the honest review-frequency answer, so the Bot knows how much unsupervised time its output will sit in. |
| The business | Supplies vocabulary. Everything the Bot writes inherits these words. |
| My systems, and how to reach them | The connection ladder, populated with the respondent's actual system names and split into connector-first and browser-only. |
| Phase 0 | Before the roster, with an explicit instruction to stop and report if any item is unmet. |
| What I want you to own, in this order | The scored sequence, with each job's frequency, verifiability, reversibility and reach stated inline so the Bot can see the reasoning rather than just the rank. Phase 4 items carry "you may prepare, you may never complete." |
| My boundaries | Green, amber and red, with the red list in the respondent's own words plus four non-negotiable additions. Ends with the reversibility tiebreaker. |
| Untrusted content | Its own section rather than a line in the boundaries, because it is the instruction most likely to matter in a year and the one most likely to be skimmed. |
| How to write as me | Voice extraction instructions, the never list, and the anti-slop file that accumulates corrections. |
| How to report to me | The evidence contract and the silent-failure guard, which is the single line that makes a broken routine look different from a quiet one. |
| Pace | Four-week handover, scaled to the setup time actually available. |
| Conditional blocks | Cost discipline, data boundary, client separation and software work, each appearing only when its trigger fired. |
| Four things about this product | Last, and deliberately repeated from the opening: no dry run, no security boundary between Bots, approvals do not reverse, the far end sees you. |
Practice
How to change it without breaking it, and what it does not do well.
7.1Adapting it
Adding a question. Append an object to the question bank with a section id, a unique id, a type, and whether it is required. Rendering, progress and validation pick it up with no further work. If the answer should appear in the brief, add the line to the brief builder; if it should change the order, it belongs in the scoring instead, which is a much bigger decision.
Adding a module. Add the section, add its questions, and add one gate function. Gates receive the whole answer set and return a boolean, so a module can depend on any combination of earlier answers.
Changing the weights. The four constant tables are the only place scoring lives. Raising a frequency weight changes rank inside phases; changing a verifiability or reversibility value changes which phase jobs land in, which is a change to the instrument's position rather than its calibration. Do that deliberately.
Adapting to a domain. The task seed suggestions are keyed to business shape and are the cheapest thing to localise. Adding a shape means adding its seed list and, usually, one module.
7.2Known limits
- It trusts the respondent's scoring. Somebody who marks everything as checkable in thirty seconds gets a fast, wrong sequence. The behavioural questions catch some of this and the divergence rules catch more, but the instrument has no independent evidence and does not pretend to.
- It has no view on whether a job should exist. A task that is pure waste scores the same as one that is essential. Automating the first is worse than not automating anything, and the instrument will not tell you which is which.
- Duration is self-reported and usually understated. People estimate from the good runs. The hours figures should be read as a lower bound.
- It cannot see interdependence. Jobs are scored independently, so two tasks that only make sense together may be split across phases. A human reading the sequence should look for that.
- Frequency weights assume a stable week. Strongly seasonal work is scored at its average, which understates the peak and overstates the trough. The seasonality question is collected but does not currently enter the formula.
- It is a snapshot of a product that changes weekly. Every constraint the brief cites was verified on 2 September 2026. The four product facts in the closing section are the ones most likely to move, and the ones whose movement would most change the output.
COMPANION TO THE GROK BOT INTAKE · Built on The Grok Bot Operator's Manual, verified against SpaceXAI documentation and Cursor billing pages on 2 September 2026. The question bank and scoring tables in this document are generated from the tool's own source, so they cannot drift apart. Nothing here is financial, legal or security advice.