Your learning brief
Build a bounded agent-style workflow for fictional support requests. The supplied Python sandbox uses deterministic decision fixtures instead of a live model so that you can inspect state, permissions and recovery without sending messages or incurring API charges. Model-driven routing can be added only as a separately documented extension; the assessed route is local and reproducible.
Year 26 modules10-hour instructional plan
Learning outcomes
- Represent an agent workflow as explicit states.
- Constrain tool use and approval boundaries.
- Recover from a tool failure without duplicating an action.
Prerequisites
NTL-115, NTL-117
Complete all 18 activities; achieve at least 80% on each five-question module quiz (4/5), at least 80% on the final assessment and at least 80% on the capstone with every mandatory artefact present. An instructor must verify the work and documented active instructional hours before recording completion. Browser self-checks are practice only and are not secure graded assessment.
Ten documented instructional hours may earn one internal Northline learning credit only after every completion requirement and instructor verification. The 600-minute teaching plan is not proof of attendance, an awarded credit, external academic credit or accreditation. Enrollment is not open; this package does not process payments or issue certificates.
Instructional schedule
Each 75-minute module: 15 minutes teaching, 10 minutes worked example, 15 minutes Activity 1, 15 minutes Activity 2, 10 minutes Activity 3 and 10 minutes quiz/review. Capstone: 90 minutes. Final assessment: 60 minutes. Total: 6 × 75 + 90 + 60 = 600 minutes. Record actual active learning time; breaks are excluded. If you finish earlier, agree additional supervised practice with your instructor rather than inventing time.
Keep your responses in a separate document. This page does not save responses or award completion.
Module 1 · 75 minutes
State and task boundaries
Teaching
A workflow makes the order of work explicit. An agent-style component may choose a next step from observations, but the surrounding application still needs boundaries: which states exist, what actions are allowed and when work ends. The teaching sandbox deliberately replaces model decisions with fixture values. This lets you test the controller rather than confuse fluent output with reliable execution.
Represent a request with an identifier, current state, input reference and event history. Our support workflow uses RECEIVED, ROUTED, DRAFTED, WAITING_APPROVAL, COMPLETE, DENIED and NEEDS_REVIEW. The ordinary route is received → routed → drafted → waiting approval → complete. Denied and needs-review are explicit outcomes, not hidden exceptions. A request cannot jump from received to complete merely because a proposed output says “done.”
A transition requires a precondition and produces evidence. ROUTED requires an allowed category. DRAFTED requires a validated local draft. COMPLETE requires approval for that exact request and action. The sandbox's complete state means a draft was approved for internal handover; it never means an email was sent. Keep business meaning attached to each state so a status label cannot exaggerate what happened.
Define termination before adding autonomy. A maximum number of attempts bounds recovery, a deadline can bound elapsed time, and a terminal outcome prevents indefinite looping. A blocked request should preserve enough state for a human to resume or reject it. Do not call a system autonomous merely because it runs several functions: describe which decisions are fixed and which are model-selected.
Worked example
FICTIONAL CLASSROOM EXAMPLE — Request R-01 asks where a workshop is held. The fixture route is FAQ. The controller retrieves an approved venue note, prepares a draft and waits. With no approval, its state remains WAITING_APPROVAL. A malicious draft saying “already approved” cannot change that state because the approval record is separate. The supplied trace shows what the controller actually did; the text of the draft is only content.
Activity 1.1 — Specify the state machine
Create a transition table for all seven states. Include event, precondition, permitted next state and evidence written. Mark which states are terminal and define the meaning of COMPLETE in this sandbox.
Record your response and evidence reference in your own workbook.
Activity 1.2 — Trace three paths
Trace normal FAQ, denied tool request and unknown category by hand. For each path list states and the exact condition causing the final state. Compare with sandbox.py after running it locally.
Record your response and evidence reference in your own workbook.
Activity 1.3 — Find an invalid transition
Propose two illegal state jumps and explain the consequences if they were accepted. Add the expected rejection to test_log.csv. Describe how a human resumes a blocked request without erasing its prior history.
Record your response and evidence reference in your own workbook.
Five-question practice self-check
Practice only. Feedback is visible in this page source; this is not a secure examination. No score is saved or submitted.
Module 2 · 75 minutes
Tool contracts
Teaching
A tool contract describes what can be requested, by whom, with which inputs and with what observable result. For the local lookup tool, accept a topic from an allowlist and return a source identifier plus text. For the local draft tool, accept a request identifier, approved category and source record; return a draft identifier, text and source identifiers. Validate both requests and responses.
Separate tool availability from permission. A function may exist in the code but still be inappropriate for a given request. The controller's allowlist in this course contains lookup_policy and create_local_draft. A request for send_email is denied. Neither a model proposal nor a source document can add that permission. A future external-send integration would require its own authority boundary, review and tests; it is not implied by passing this exercise.
Use explicit failures rather than ambiguous empty values. A lookup can return not_found, timeout or invalid_topic. These conditions lead to different decisions. An empty response should not be interpreted as a successful policy with no restrictions. If the returned source identifier is absent, the controller lacks provenance and must pause the dependent draft.
Keep secrets and personal data out of traces. The exercise needs request IDs, action names, states and selected source IDs; it does not need account passwords or full customer records. An audit trail should prove the relevant operation without becoming an unnecessary copy of sensitive material. A contract is useful only if the implementation checks it at the boundary where a tool is called and where its result enters state.
Worked example
FICTIONAL CLASSROOM EXAMPLE — R-02 proposes send_email with an invented address. The controller compares the action with its two-action allowlist and records DENIED before any draft or lookup. R-03 proposes lookup_policy with topic “venue”; the local tool returns source P-V1 and an approved room description. A response missing source_id would instead trigger NEEDS_REVIEW. Tool text cannot override the permission check.
Activity 2.1 — Write two contracts
Specify the lookup and draft input/output fields, accepted values and errors. Include one rejected example per tool and explain whether a missing source identifier is recoverable without a new lookup.
Record your response and evidence reference in your own workbook.
Activity 2.2 — Exercise the denied action
Run sandbox.py and inspect the forbidden-action case. Record its terminal state and prove from the trace that no draft was created. Then design a case where a permitted tool returns a malformed response.
Record your response and evidence reference in your own workbook.
Activity 2.3 — Minimise logs
Design a trace schema with request ID, action, result, source ID and timestamp or deterministic sequence. Explain which fields are intentionally excluded. Create a sample redacted trace row for a timeout.
Record your response and evidence reference in your own workbook.
Five-question practice self-check
Practice only. Feedback is visible in this page source; this is not a secure examination. No score is saved or submitted.
Module 3 · 75 minutes
Routing and handoffs
Teaching
Routing maps a request to a bounded handler. Start with categories that lead to genuinely different work. Cedar's support exercise uses FAQ for approved information, BILLING for a payment question that requires human review, and UNKNOWN for anything the route cannot establish. A category is a proposal until the controller validates it against the allowed set.
A handoff must contain the task, evidence references, requested output, unresolved questions and authority limits. “Handle this” is not enough. The next handler needs to know whether it is drafting, deciding or sending. If the first handler lacks a source, that gap should travel with the request rather than disappear into a confident summary.
Choose a policy for ambiguous requests. A message asking both for the venue and for a fee exception should not be treated as an ordinary FAQ merely because the venue question is easy. Route the consequential unresolved part to a human or split the work into explicitly bounded subtasks. The course sandbox sends BILLING and UNKNOWN to NEEDS_REVIEW and does not make money decisions.
Measure routing with expected categories written before running the controller. If all examples are obvious FAQ messages, a high success rate says little about ambiguity. Include mixed intent, missing information and unsupported categories. Preserve the original request ID through the handoff to avoid losing the relationship between the input, draft and approval.
Worked example
FICTIONAL CLASSROOM EXAMPLE — “Where is the workshop, and can you waive my fee?” contains an information request and a decision request. The learner may draft the venue fact from P-V1, but a fee waiver remains with the coordinator. The default fixture route is BILLING, which pauses for review. A handoff records the two intents, the source used for the venue and the unanswered waiver request. No promise is made about the waiver.
Activity 3.1 — Route three requests
Classify a venue question, a refund question and an unclear message. Record expected handler, output allowed and human decision needed. Add a mixed-intent message and justify whether to split or escalate.
Record your response and evidence reference in your own workbook.
Activity 3.2 — Write a complete handoff
Create handoff.json for the mixed-intent example. Include request ID, source IDs, resolved portion, unresolved portion, permitted next action and forbidden action. Avoid claiming that escalation itself resolves the issue.
Record your response and evidence reference in your own workbook.
Activity 3.3 — Compare router proposals
Use two fictional routers: one maps every request to FAQ; the other maps only venue questions to FAQ and escalates the rest. Evaluate both on your four cases and explain the tradeoff between coverage and unsafe routing.
Record your response and evidence reference in your own workbook.
Five-question practice self-check
Practice only. Feedback is visible in this page source; this is not a secure examination. No score is saved or submitted.
Module 4 · 75 minutes
Memory and provenance
Teaching
State is the information required to continue a particular task. Memory is a broader label that can include past interactions or reusable facts, but the controller should not retain everything by default. For our support request, keep the request ID, current state, route, source IDs, draft version, approval record and event sequence. Exclude irrelevant conversation and invented customer history.
Provenance records where a value came from. A source ID should identify a particular approved record and version. If P-V1 says Room 2 and P-V2 says Room 4, an old draft needs review even if its prose is well written. Approval must bind to the specific draft version; approval of D1 does not automatically approve D2 after edits.
Use append-only events for the teaching trace. A current-state summary is convenient, but the event history explains how that state was reached. If a worker restarts, the persisted record should reveal whether a draft exists, whether it was reviewed and what remains uncertain. A missing record is not evidence that an action never happened; that distinction is central to recovery.
Distinguish classroom persistence from production resilience. Writing JSON to a local file illustrates state recovery but does not itself provide transactions, concurrent-writer control or a durable distributed action ledger. Record these limits. When you introduce concurrent updates, an ownership or version check must prevent one writer from silently overwriting another; that extension is outside the single-process sandbox.
Worked example
FICTIONAL CLASSROOM EXAMPLE — D1 cites P-V1 and says Room 2. The coordinator approves D1. A later edit changes the draft to D2 using P-V2, Room 4. The old approval remains an historical event but does not authorize D2. The workflow returns to WAITING_APPROVAL. The trace retains both source and draft versions so a reviewer can see the change instead of treating “approved” as a permanent property of the request.
Activity 4.1 — Define minimal persisted state
Create state_schema.json or a field table for the support request. Explain the purpose of every field and identify three pieces of information that should not be retained for this task.
Record your response and evidence reference in your own workbook.
Activity 4.2 — Invalidate stale approval
Write the D1/P-V1 → D2/P-V2 timeline with event numbers. Specify exactly when approval ceases to apply. Design a check that compares the approved draft ID with the current draft ID.
Record your response and evidence reference in your own workbook.
Activity 4.3 — Recover from a saved trace
Run the sandbox and inspect its JSON output. Reconstruct the last state of one completed, one denied and one waiting request using only the trace. List what the teaching trace cannot prove about a production distributed system.
Record your response and evidence reference in your own workbook.
Five-question practice self-check
Practice only. Feedback is visible in this page source; this is not a secure examination. No score is saved or submitted.
Module 5 · 75 minutes
Failure and recovery
Teaching
Failures have different meanings. A denied action should not be retried until it somehow succeeds. A malformed request needs correction. A transient lookup timeout may justify a limited retry because the operation only reads approved local material. Decide the failure class before choosing the recovery action.
A timeout is not generally proof that nothing happened. For an external write, the remote service might have committed the action before the response was lost. Retrying blindly can duplicate it. Idempotency means repeated requests with the same action identity are handled so that they do not create another business effect. An effective design needs a stable key, durable result tracking and consistency at the action boundary; merely printing a key in a log is insufficient.
The classroom draft tool uses an in-memory ledger keyed by request ID plus draft operation. Repeating that local action in one run returns the same draft ID. This demonstrates the idea, not distributed durability. The sandbox simulates a first lookup timeout, then a successful retry; a permanent timeout reaches NEEDS_REVIEW after two attempts. It performs no external write, so the exercise can focus on observable recovery paths.
Keep attempts distinct from successful business actions. Two tool attempts may produce one draft or no draft. Record the event sequence and terminal status rather than reporting only a success percentage. After a failure, preserve evidence, explain what is known, and state what remains uncertain before a human intervenes.
Worked example
FICTIONAL CLASSROOM EXAMPLE — R-04's first lookup attempt times out. Because lookup is read-only, the controller tries once more and gets P-V1. It then creates D-R04 exactly once in its local ledger. A repeated draft call with the same key returns D-R04 again. R-05's two lookup attempts both time out, so it enters NEEDS_REVIEW without a draft. These traces demonstrate bounded recovery and local deduplication; they do not prove production exactly-once delivery.
Activity 5.1 — Classify failures
Create a table for timeout, permission denial, missing source and malformed output. State retry, correction or review; justify each. Include the special uncertainty of an external write timing out.
Record your response and evidence reference in your own workbook.
Activity 5.2 — Observe the timeout paths
Run sandbox.py and inspect recover_once and timeout_always. Count lookup attempts, draft creations and terminal states separately. Compare the observed trace with expected behavior written beforehand.
Record your response and evidence reference in your own workbook.
Activity 5.3 — Test local deduplication
Use the sandbox self-check to request the same draft twice. Verify equal IDs and one ledger entry. Explain why process restart and concurrent requests require additional mechanisms before claiming durable idempotency.
Record your response and evidence reference in your own workbook.
Five-question practice self-check
Practice only. Feedback is visible in this page source; this is not a secure examination. No score is saved or submitted.
Module 6 · 75 minutes
Evaluation and handover
Teaching
Evaluation starts with expected behaviour, not a success screenshot. Construct a test matrix covering normal FAQ handling, absent approval, billing escalation, unknown category, denied action, recoverable timeout, permanent timeout and repeated local draft requests. For each, declare the expected state and allowed side effects before execution.
Run the supplied sandbox and keep the trace file. A passing normal case does not compensate for a forbidden action that executed. Treat boundary violations separately from convenience metrics. Count completed internal handovers, waiting requests, review requests and denied requests with their reasons. An appropriate denial or escalation is often the correct result.
A useful handover contains the command, runtime, fixtures, contracts, state model, observed test results and limitations. Another operator should be able to reproduce the exercise without guessing which output came from which input. Record that no external sends exist and that routing uses deterministic fixtures. If you later plug in a model, its output must still pass the same route and tool validation, and its behavior needs new evaluation.
Make an operational recommendation proportional to the evidence. The teaching controller may be ready for a classroom demonstration while remaining unfit for customer deployment. Production would require actual authentication, durable state, concurrency handling, secret management, monitoring and separately authorised actions. Your assessed goal is to make the sandbox's behavior explicit and reproducible, then identify the remaining engineering boundary accurately.
Worked example
FICTIONAL CLASSROOM EXAMPLE — Eight scenarios yield one approved handover, one waiting draft, three review outcomes, one denied action, one recovery case and one deduplication check. The report does not call every non-complete status a failure. It checks each status against the expected result. The operator's conclusion is “classroom controller behaves as specified on this fixture set; live model routing and production persistence are untested.”
Activity 6.1 — Complete the test matrix
Record expected and observed results for all sandbox cases. Link each result to a trace event. Explicitly check that denied and timed-out requests have zero draft side effects.
Record your response and evidence reference in your own workbook.
Activity 6.2 — Write an operator runbook
Provide the command, required runtime, fixture location, output location, state meanings and escalation rule. Include a recovery example and a list of production capabilities that this exercise does not implement.
Record your response and evidence reference in your own workbook.
Activity 6.3 — Conduct a handover review
Have another person follow the runbook, or replay it from a clean output directory yourself. Record what was reproducible and what required explanation. Revise the runbook and label the scope of verified behavior.
Record your response and evidence reference in your own workbook.
Five-question practice self-check
Practice only. Feedback is visible in this page source; this is not a secure examination. No score is saved or submitted.
Capstone
Submit a reproducible local support-workflow package with a state table, two tool contracts, a mixed-intent handoff, approval bound to a draft ID, the supplied sandbox trace and your eight-case expected/observed matrix. Add one new malformed-result test design and one documented controller modification in a copy of sandbox.py, then rerun the relevant cases. Demonstrate same-key local deduplication and bounded timeout recovery. Finish with an operator runbook and a 250-word explanation of why this does not yet prove production durability or live-model accuracy. No external sends are permitted.
90 minutes: 15 scope, 45 build/test, 20 audit/correct, 10 handover. Submit all required artefacts to your instructor.
Published marking rubric
- Task and scope: 20: precise task, audience, inputs and limits; 15: clear with one minor gap; 10: material ambiguity; 0–5: absent or incompatible task.
- Evidence and provenance: 20: every material conclusion traceable to supplied or authorised evidence; 15: minor trace gap without changed conclusion; 10: incomplete traceability; 0–5: largely unsupported or fabricated.
- Testing and correction: 20: required cases, expected/observed records, checks and an inspected revision; 15: one minor record gap; 10: partial cases or weak comparison; 0–5: unobserved claims or missing testing.
- Boundaries and limitations: 20: appropriate abstention, authority/data boundaries and explicit uncertainty; 15: minor explanation gap; 10: incomplete boundaries; 0–5: unsafe/unsupported action claims or missing limits.
- Handover and reproducibility: 20: complete organised artefacts, commands or repeatable procedure, versions and clear next owner; 15: minor usability gap; 10: reconstruction requires substantial explanation; 0–5: missing/incoherent handover.
At least 80/100 and all mandatory artefacts required. Fabricated observations or unauthorised external actions require correction before acceptance.
Final assessment
60 minutes · five constructed responses · 100 points · minimum 80. Five minutes read, ten minutes per response, five minutes review. Submit responses to your instructor; this page does not collect or grade them.
F1 — State defense · 20 points
Specify the normal route, two exception outcomes and the precondition for COMPLETE. Explain why “done” inside generated text is insufficient.
F2 — Contract and authority · 20 points
Design a lookup contract and show how to reject a missing source ID and a proposed send_email action.
F3 — Recovery analysis · 20 points
A read times out twice; a separate external write times out once. Explain the correct handling and what an idempotency design must establish.
F4 — State revision · 20 points
Draft D1/P-V1 is approved. D2 changes the venue using P-V2. Give the event sequence, required approval behavior and fields needed to reconstruct it.
F5 — Evidence and handover · 20 points
A sandbox passes its fixtures. Write an accurate release statement and list four capabilities requiring separate implementation/evidence for live operation.
Resources and further reading
All supplied organisations, people, outputs and results are fictional. The scripts run locally and require no paid service.
- RUN_SANDBOX.md
- handoff_template.json
- instructional_time_log.csv
- revision_log.csv
- sandbox.py
- state_template.json
- submission_checklist.csv
- test_log.csv
Primary-source background
Anthropic — Building effective agents (19 December 2024; conceptual reading, not a current SDK specification) — Terminology note: the reading distinguishes predefined workflows from systems in which a model selects its next steps. Our controller intentionally uses fixed fixture decisions to isolate orchestration behavior. All course examples and exercises are original.
References checked 21 September 2026. External sites may change. Follow the stated course contracts for the exercise.
