Agentic AI — Day 5
What We're Going to Cover
Day 4 gave the Weekly Check-In Agent hands: it could read FitPath data and act on it, with an approval gate for risky users. But it was still one agent doing everything — reading, judging risk, and acting, all inside a single instruction block.
Today you split that agent into a supervisor and three specialized sub-agents. You'll design the delegation logic, wire it in Fleet, configure a human checkpoint, and then prove it behaves differently for a clean user than for a flagged one.
By the end you'll have a working multi-agent system and be able to answer, concretely: where did specialization actually earn its complexity, and where did it just add overhead?
Agentic AI — Day 5
Context
sam_w— clean profile, no injuries, no clearance issues.jordan_e— knee pain reported, no medical clearance on file.
Same MCP server as Day 4 (Postgres-backed, hosted on Render), exposing the same 8 tools:
| Read | Write |
|---|---|
get_user_profile | log_workout |
get_workout_history | send_notification |
get_equipment_inventory | escalate_to_trainer |
get_action_log | update_equipment_status |
A Fleet constraint that shapes this whole exercise
Sub-agents can't have individual tools removed from them — every sub-agent can technically see and call all 8 tools. The only per-tool control you have is Auto (runs freely) or Ask (needs human clearance first), and that's set per sub-agent, per tool. This means "restrict this sub-agent to only its own tools" isn't something the platform will do for you — you have to get the same result through instructions plus universal write-gating instead. That's covered in Step 2.
Setup note
The server identifies you by the X-Participant-Id header on your MCP connection — the same one you configured in Day 4, not a suffixed user ID. If you're building fresh sub-agents rather than reusing a Day 4 connection, make sure each one is wired to a connection carrying that same header, or your reads/writes will land in the wrong participant's data.
The task this system handles: a weekly check-in for a given user — look at what happened this week, decide whether it's safe to log/notify/adjust automatically, and act (or escalate) accordingly.
The difference from Day 4: instead of one agent doing all of that, you're building three specialists plus a supervisor that coordinates them. This is the Supervisor pattern — not a router, because the orchestrator isn't just classifying and forwarding a request to one path; it's combining input from multiple agents before deciding what to do next.
On agent count
These three sub-agents are built as nested Add subagent entries under one orchestrator card, not as three separate top-level Fleet agents — so this should fit inside a free-tier workspace even with the 1-agent cap. If you hit a limit while building, flag it immediately rather than working around it silently.
Agentic AI — Day 5
Step 1 — Design Before You Build ⏱ 10 minutes
| Question | Your answer |
|---|---|
| What are the three sub-agents and what does each own? | |
| Which of the 8 tools does each sub-agent actually need to call — and how will you stop it calling the rest, given every sub-agent can technically see all 8? | |
| What does the supervisor do with the outputs of the two analysis agents before deciding whether to act? | |
| What conditions should cause the supervisor to withhold action and require human approval? |
If you get stuck on the sub-agent split, here's the shape (fill in the reasoning yourself — don't just copy this):
- Progress Analyst — reads what actually happened this week. Should never write, even though it technically can.
- Safety Reviewer — reads the user's risk profile and decides if this user is safe to act on autonomously. Should never write, even though it technically can.
- Action Coordinator — the sub-agent meant to actually execute writes. Executes only what the supervisor tells it to, and only after any required approval clears.
Agentic AI — Day 5
Step 2 — Build the Sub-Agents and Configure the Human Checkpoint ⏱ 30 minutes
Add subagent, same as the panel you've already seen). For each one, write instructions that specify:- Scope — what this agent does and does not do (be explicit about the "does not" — that's what keeps it from overreaching).
- Tool control — since you can't remove tools from a sub-agent, set every tool's Auto/Ask state deliberately rather than leaving defaults. For all three sub-agents: leave the four read tools (
get_user_profile,get_workout_history,get_action_log,get_equipment_inventory) on Auto — an off-scope read is wasteful, not unsafe. Set all four write tools (log_workout,send_notification,escalate_to_trainer,update_equipment_status) to Ask on every sub-agent, not just the Action Coordinator. For the Progress Analyst and Safety Reviewer this should be a safety net that never fires — if their instructions are correct, they never attempt a write, so the Ask gate is just insurance against an instruction failure, not something you expect to see triggered in normal testing. - Output — what it should hand back to the supervisor (a short structured summary, not a wall of prose — the supervisor has to read this).
A design choice to make deliberately, not by default
In Day 4, both FitPath source documents (the exercise/programming guide and the safety guidelines) were uploaded into the agent's Memory. You can carry that forward into the Safety Reviewer here — giving it the actual written safety policy to reason against, instead of just raw profile fields. That's closer to how this would really work, but it also means the Safety Reviewer's judgment is partly grounded in a document you didn't write today. Decide which you want for this exercise and note it in your placeholder below — both are defensible, but you should be able to say which one you chose and why.
Starter skeleton for the Safety Reviewer — the other two follow the same shape, write them yourself:
Safety Reviewer scope
?
state exactly what data this agent reviews (profile fields only, or something broader) and why that's enough for a safety judgment.
Memory grounding
?
state whether this agent has the FitPath safety guidelines in Memory, and if so, instruct it to apply that document's rules rather than inventing its own.
Reads it doesn't need
?
explain why the Safety Reviewer doesn't need workout history or the action log to do its job.
Evidence for a flagged status
?
specify what information must accompany a FLAGGED result so a human reviewer can act on it without pulling more data.
For every sub-agent, configure tool control now, while you're already in each panel
Fleet's tool control is Auto / Ask, not a separate approval toggle — setting a tool to Ask is what requires a human to clear it before it runs; that's the same mechanic as Day 4's "Set the Dial," just naming it correctly this time. The gate is per-tool, not conditional on what a sub-agent found — that's a real constraint, not a bug in the exercise: it means the content of the approval request has to do work the trigger itself can't.
- Set all four write tools to Ask, on the Progress Analyst and Safety Reviewer as well as the Action Coordinator. Leave the four read tools on Auto for all three.
- For the Action Coordinator specifically: write into its instructions what an escalated request must contain: the user ID, the flagged reason it was given, and what action was being considered — enough context that whoever's reviewing can make the right call without opening a second tab. You don't know exactly what the orchestrator will hand it yet (that's Step 3), so write this generically — "include whatever flagged reason the supervisor gives you," not "include the Safety Reviewer's exact output." The interface between them gets nailed down next.
This is the design point around Information Shown when designing human checkpoints. A checkpoint that shows nothing useful isn't a checkpoint, it's a rubber stamp.
📋 Save your work — Sub-agent instructions
Paste the final ROLE / OUTPUT text you actually used for each sub-agent, after you've settled on it. Not a summary — the real text, so you (or the facilitator) can see exactly what the agent was told.
Progress Analyst — final instructions
Write tools set to Ask? Y / N
Safety Reviewer — final instructions
Did you include the Memory-grounded safety document? Y / N — why
Write tools set to Ask? Y / N
Action Coordinator — final instructions (including its approval-request behavior)
Write tools set to Ask? Y / N
Agentic AI — Day 5
Step 3 — Write the Supervisor's Delegation and Aggregation Logic ⏱ 15 minutes
Instructions field, in prose. Get this wrong and nothing downstream saves you.Your orchestrator's instructions need to cover:
- When it calls the Progress Analyst and Safety Reviewer (should this always run both, every time? Why?)
- How it reconciles their outputs if they disagree or one is inconclusive — don't skip this; write down what "disagree" would even look like for these two agents
- What it tells the Action Coordinator to do, and under what condition it does not authorize the Action Coordinator to run at all
Starter skeleton:
Delegation input
?
state exactly what you send to each analysis sub-agent, and whether both are called before you know anything else.
Aggregation logic
?
state precisely how the two outputs are combined into one decision, including how a Safety Reviewer FLAGGED overrides the Progress Analyst's findings.
Action Coordinator handoff
?
state what you authorize on a CLEAR outcome and what you withhold on a FLAGGED outcome.
📋 Save your work — Orchestrator instructions
Paste the final ROLE / PROCESS text you actually used.
Agentic AI — Day 5
Step 4 — Test Both Users ⏱ 15 minutes
jordan_e to get the right outcome, that's a sign the orchestration logic isn't doing its job.Capture the raw output from each agent as it happens — not just your read of what happened. These placeholders are for the actual text each agent produced, before any interpretation.
Run 1 — sam_w (expect: completes end-to-end, no approval needed)
📋 Progress Analyst raw output
📋 Safety Reviewer raw output
📋 Orchestrator's decision (what it told the Action Coordinator)
📋 get_action_log raw result (proof the action actually happened)
Run 2 — jordan_e (expect: stops at an Ask gate)
📋 Progress Analyst raw output
📋 Safety Reviewer raw output
📋 The pending Ask request, exact text as it appears in Fleet's Inbox
Comparison — after both runs
sam_w | jordan_e | |
|---|---|---|
| Behaved as expected? (Y/N) | ||
| If N, which sub-agent or instruction caused the gap? |
If either run doesn't produce the expected result, don't just re-run it — look at the trace and figure out which sub-agent's instructions caused it, or whether an Ask gate fired that shouldn't have (or didn't fire when it should have). The comparison table above is for synthesis once you've already captured the raw outputs — not a substitute for capturing them.
Agentic AI — Day 5
Success Criteria
Agentic AI — Day 5
If You Finish Early
This one's a unit test of the orchestrator's aggregation logic, not another full pipeline run — you don't have write access to sam_w or jordan_e's underlying records, so you can't manufacture a real disagreement between the Progress Analyst and Safety Reviewer through the data itself. Instead, test the orchestrator's reasoning directly, with a controlled hypothetical.
Send the orchestrator something like this — note that it deliberately tells the orchestrator not to call its sub-agents, since the point here is to isolate its decision logic, not run the full system:
Watch what comes back. Does the orchestrator notice the note buried in the Progress Analyst's report and treat it as a reason to escalate, even though the Safety Reviewer said CLEAR? Or does it just follow the Safety Reviewer's status and authorize logging, because its instructions never told it to look for risk language outside the Safety Reviewer's own output?
If it's the latter — which is the likely outcome, since the current orchestrator instructions only check the Safety Reviewer's status and never inspect the Progress Analyst's content for risk signals — fix the instructions so a CLEAR from the Safety Reviewer isn't automatically the end of the check. Something like: even when Safety Reviewer returns CLEAR, scan the Progress Analyst's notes for language suggesting pain, injury, or reluctance (e.g. "pushed through," "hurt," "sore"); if present, treat it as equivalent to a FLAGGED result and route to the withheld path instead of silently proceeding.
Re-send the same hypothetical after your fix and compare.
📋 Extension — before/after
Raw orchestrator output before your fix (from the hypothetical prompt above)
Your full orchestrator instructions after the fix (the complete ROLE / PROCESS text, not just the new step)
Raw orchestrator output after your fix (same hypothetical prompt, re-sent)
Full failure-injection version of this is available after-hours if you want to go further — ask for it.
