Spec-driven development illustrated by a delicate web built within fixed anchors across a pastel sandstone ravine

Spec-driven development starts paying off when two people can read the same request and build different agents. “Handle client intake” might mean collecting a name, creating a CRM record, sending an email, or deciding whether someone qualifies for a service.

I looked through the current spec workflows and agent evaluation guidance with a more practical question: what should a consultant agree with a client before connecting the agent to real tools?

My recommendation is a short, versioned behavior spec with concrete examples. Use it to agree on the job, build the agent, and decide whether a later change is acceptable; keep the implementation evidence beside it.

What spec-driven development means for a client agent

A spec is a written agreement about observable behavior. It describes who the agent serves, what information it can use, which actions it can take, and what must be true when a task finishes.

The prompt is one implementation of that agreement. Permissions, knowledge sources, tool validation, and review rules implement other parts, so putting the entire agreement in a prompt does not finish the job.

Birgitta Böckeler’s October 2025 exploration of spec-driven development separates spec-first, spec-anchored, and spec-as-source approaches. The difference is whether you write a spec initially, maintain it afterward, or treat it as the primary source from which implementation is generated.

For a client agent, I would usually choose spec-anchored: keep the agreed behavior current while separately maintaining the configured agent and integration code. That is a proposed working method, not a claim that a Markdown document automatically compiles into a safe agent.

Our guide to writing agent instructions covers the prompt itself. This guide covers the agreement that tells you which instructions and controls you need.

Which jobs deserve a written spec?

Write one when the agent will interact with customers, change business records, use private information, or survive handoff to another person. Those situations benefit from an explicit answer to “who agreed this was allowed?”

A disposable brainstorming assistant can start with a clear prompt and a few checks. I would not require a document package for changing an ice breaker or fixing a typo.

Increase the detail with the consequences of misunderstanding. A draft-only research assistant needs source rules and output examples; an assistant that submits orders also needs spending authority, approval conditions, duplicate protection, and confirmation evidence.

Keep the first spec small enough for the client to read. Put extensive tool schemas, source inventories, and evaluation cases in linked appendices instead of hiding the main decision in a wall of text.

There is no useful page-count contest here. The right stopping point is when the client, builder, and reviewer can independently explain the same boundaries and recognize the same successful outcome.

Start with the outcome, owner, and non-goals

Write the job as an outcome for a specific user: “Help a new prospect submit complete intake information for a human to review.” That is easier to evaluate than “Be a helpful sales agent.”

Name the business owner who can resolve policy questions and the operator who can pause the deployment. They may be the same person, but neither role should be an unnamed future teammate.

Then list non-goals. In our proposed intake example, the assistant does not quote custom prices, promise project dates, reject prospects, or send outbound messages.

Non-goals prevent a reasonable-sounding interpretation from quietly expanding the project. If someone later requests automatic qualification, that becomes a visible scope change with its own evidence.

State the handoff too: the agent produces a reviewable intake package, and the account owner decides what happens next. A system with an undefined final owner can appear finished while leaving the actual work unattended.

Spec-driven development lifecycle from agreed behavior through implementation, evidence, and approved changes

A one-page starting template for spec-driven development

This is the template I would bring to a client conversation. Copy the editable agent spec template, replace the brackets, and attach examples where a sentence could mean more than one thing.

Agent: [name] | Spec version: [version] | Owner: [person]
Job: [user + useful outcome]
Non-goals: [work excluded from this release]
Inputs: [required fields, optional fields, trusted identity]
Sources: [approved source, policy owner, freshness rule]
Outputs: [format, destination, completion evidence]
Tools: [read / draft / write permissions and preconditions]
Approval: [who approves what, expiry, changed-payload rule]
Fallbacks: [missing data, conflict, failure, unknown outcome]
Budgets: [time, attempts, spend and enforced stop behavior]
Acceptance: [requirement IDs + examples + release blockers]
Operations: [logs, pause method, recovery and review owner]
Sign-off: [version, approver, date, unresolved questions]

The opening page is an index to decisions, not a substitute for them. “Use the CRM safely” is still unresolved until the tool appendix names the allowed fields and the cases demonstrate what happens when a lookup fails.

Leave unknowns visible. Write “approval expiry: awaiting the account owner’s decision” rather than inventing a deadline that looks authoritative because it appears in a template.

Worked example: an intake assistant that stops at review

Consider a fictional consulting agency receiving requests through a website. The proposed assistant helps visitors describe a project, checks the approved service guide, and prepares an intake package.

Its required inputs are a contact name, verified contact route, project description, and consent to submit the information. Budget range is optional, so a visitor can complete intake without disclosing one.

The output is a structured draft with contact details, the visitor’s own description, missing optional information, and relevant source references. The interface clearly labels it ready for review, rather than promising that an engagement has been accepted.

For the initial release, I would keep CRM writes behind a human approval step and disable outbound email. This reduces the number of external consequences while the agency checks whether the intake package is actually useful.

The completion condition is precise: either a complete package is ready for authorized review, or the visitor receives a clear explanation of what is missing and how to reach a person. A friendly farewell is not sufficient evidence that intake finished.

This example is a design proposal throughout. It does not describe a customer deployment or claim a measured reduction in intake time.

Turn vague requirements into observable examples

“The agent should be accurate” gives a reviewer very little to work with. “When the service guide does not contain a requested price, the agent says it cannot quote that price and offers a human handoff” identifies an observable boundary.

I use a simple Given, When, Then structure to expose missing decisions. Cucumber’s Gherkin reference defines these as starting context, an event, and an expected outcome.

For example: Given a prospect has supplied all required fields but no budget, when they submit the draft, then the package can proceed to review with budget marked “not provided.”

A second example: Given no consent has been recorded, when the visitor asks the assistant to submit intake, then it requests consent and creates no CRM record.

A third: Given the visitor changes the contact address after reviewing the package, when they submit again, then the assistant presents the revised destination and requires the appropriate fresh confirmation.

These are plain-language requirements, not executable Cucumber tests. They still do useful work because a client can disagree with the expected outcome before you build the wrong behavior.

Specify inputs and sources before choosing a model

Separate information the user provides from identity verified by the application. A name or email typed into a chat should not automatically grant access to an existing customer’s records.

For each required input, state the response to an empty value, a malformed value, and a later correction. Decide whether the assistant can ask a follow-up question or must hand off immediately.

For knowledge, name the source owner and conflict rule. In our example, the approved service guide wins over an old uploaded brochure; if two approved guides conflict, the assistant escalates rather than selecting whichever passage sounds persuasive.

Our knowledge base setup guide explains the retrieval layer. The spec adds the policy decision about which retrieved information the agent may rely on.

If outputs enter another system, define their required fields and types. The JSON Schema getting-started guide demonstrates structural requirements, but a valid structure alone does not prove that a contact belongs to the current user.

Three action boundaries for a client agent: read approved sources, draft for review, and write after approval

Give every tool a permission and a precondition

A tool list should say more than “CRM, calendar, email.” Record which operations are available, which resource they can touch, and what must be true before each operation runs.

I would separate the intake assistant’s tools into three groups:

  • Read: retrieve the approved service guide and authorized intake state.
  • Draft: assemble the proposed intake package without creating an external record.
  • Write: submit the exact approved package to the permitted CRM destination.

The write operation needs an authorized approver, a specific payload, and a check that the approval still applies. A general “looks good” earlier in the conversation should not authorize a changed destination or newly added data.

Enforce the boundary in the integration. A prompt can explain a permission rule, but the connected tool or workflow must reject a write that lacks valid authorization.

Anthropic’s engineering guidance on effective agents emphasizes carefully documented tools, environment feedback, and stopping conditions. I would treat tool definitions as part of the implementation that must be checked against the spec.

For more detail on choosing oversight per action, see our human-in-the-loop guide. The useful question here is which consequence requires a person, rather than whether the whole agent is “autonomous.”

Write down what happens when something fails

Happy-path requirements are cheap. The decisions that protect the client usually appear when information is missing, sources conflict, or a service returns an ambiguous result.

Give each failure a user response, state transition, and owner. If the service guide cannot load, the assistant should explain that limitation, preserve the draft if appropriate, and identify the human handoff instead of improvising service details.

A write timeout deserves separate treatment. The CRM might have created the record even though the assistant never received confirmation, so “try again” can produce a duplicate.

In the proposed design, the integration uses a stable submission key and looks up the write outcome before attempting another creation. If the destination cannot support duplicate protection or reliable reconciliation, keep the action under manual control until the builder supplies an equivalent safeguard.

The status should remain outcome unknown until there is evidence. Do not translate an unclear tool result into either “submitted successfully” or “nothing happened.”

Our multi-step workflow guide explains gates and action handoffs. The spec records the business decision for each branch so the implementation has something concrete to follow.

Define evaluation cases before polishing the conversation

Ask the client for situations where a wrong response would matter: a missing consent, an unsupported service, an impatient visitor, or an existing submission. Use sanitized or synthetic inputs when the originals contain private information.

Each case needs starting state, user input, expected behavior, and evidence. For a CRM write, inspect the resulting record and authorization trail instead of grading only the final chat message.

Anthropic’s January 2026 agent evaluation guidance distinguishes the transcript from the final environment outcome and describes code-based, model-based, and human graders. That distinction is particularly useful when an agent says it completed an action.

Separate critical boundaries from overall quality. An unauthorized write should not disappear inside a high average score for friendly conversation.

For the fictional intake pilot, I would block release on any observed unauthorized write, unsupported price promise, or duplicate record. Other thresholds, such as how often clarification is needed, should be agreed with the client before running the suite.

Keep a development collection for iteration and a separate regression collection for release comparison. Our guide to building an agent evaluation set goes deeper on replay, graders, repeated attempts, and interpreting changes without score shopping.

Requirement traceability flow connecting an agreed rule to a test case, implementation control, and release evidence

Keep a requirement-to-evidence table

Assign short IDs to requirements, then link each one to its implementation and checks. This gives reviewers a practical way to ask whether an important decision has actually been covered.

RequirementImplementation controlEvaluation caseEvidence to inspect
R1: Collect required fieldsValidated intake form or stateMissing contact routeDraft remains incomplete; no write
R2: Respect approved sourcesSource selection and conflict handlingBrochure disagrees with guideAnswer follows guide or escalates
R3: Write only with approvalAuthorization check in integrationPayload changes after approvalChanged write rejected
R4: Avoid duplicate submissionSubmission key and outcome lookupTimeout after record creationOne record; reconciled status
R5: Stop within agreed budgetEnforced workflow limitsRepeated unavailable lookupStopped run with handoff

This table is illustrative, not a completed test report. An entry saying “authorization check” is an implementation task until the corresponding evidence shows the boundary works.

It also exposes orphan work. If a clever new feature has no requirement, either justify the scope change or remove it; if a requirement has no check, the release evidence is incomplete.

Resolve open questions before asking for sign-off

Use an ambiguity review: ask what two reasonable readers could interpret differently, then bring those differences to the owner. An agent can help generate questions, but it should not invent the client’s policy answers.

For intake, the questions might be whether consent covers CRM storage, who reviews submissions, what happens to unfinished drafts, and whether visitors can withdraw or correct information.

Give each open question an owner and mark whether it blocks the proposed release. Cosmetic copy can remain open while you build a prototype; an unresolved write permission cannot.

Sign-off should identify a version and its boundaries. “Approved spec v1.0 for draft intake only” is useful evidence; “approved the AI project” leaves too much room for disagreement.

After sign-off, keep a copy beside the cases and release record. If the client approved one version while the builder implemented another, the documents need reconciliation before deployment.

Translate the agreed spec into a Pickaxe agent

In Pickaxe, I would use the spec to organize configuration across Build, Knowledge Base, Actions, and Preview. The spec stays as the external agreement; I am not describing a built-in spec compiler.

The System prompt in Build states the job, non-goals, response style, and escalation rules. Knowledge Base contains the approved source material, while Actions connect only the operations needed for the agreed task.

Use Preview to try the agreed examples, then check the actual deployment and connected workflow. A good preview answer does not establish that external approval enforcement, resource isolation, or duplicate protection works.

For the draft-only intake release, the agent can collect information and prepare the package while submission remains a separate controlled step. Add the write capability only when the integration implements the agreed checks.

If clients access the agent through a Portal, configure the intended audience through the appropriate Access Groups and verify the permissions from a representative user’s session. Keep integration authorization narrower than whatever information the chat happens to contain.

Pickaxe’s Agent Builder guide covers the configuration interface. The release record should capture which model, knowledge sources, prompt, and connected actions correspond to the approved spec.

Do you need GitHub Spec Kit or Kiro?

You can start with ordinary Markdown. Choose a tool when its workflow helps your team maintain decisions and evidence, rather than because the document needs a particular filename.

The current GitHub Spec Kit quickstart, checked October 6, 2026, carries work through specify, plan, tasks, implement, and converge, with additional clarification and consistency checks available. It is useful when a coding agent builds the integration around your client agent.

Kiro’s current spec documentation describes requirements or bug analysis, design, and tasks artifacts. Its feature spec guide supports requirements-first and design-first workflows and structured requirements using EARS notation.

Those tools structure software development work. A behavioral agreement for a no-code agent can borrow their discipline without adopting their entire coding workflow.

None of these documents substitutes for a working permission check or credible release evidence. I would choose the smallest process that keeps the approved behavior, configured agent, and tests aligned.

Maintain the spec when the agent changes

A maintained spec needs a change process. When someone requests a new action, source, audience, or output destination, identify which requirements change before editing the prompt.

Use a compact change record: request, affected requirement IDs, new behavior, new risks, updated cases, approver, and release version. Attach the evidence from the candidate configuration.

For example, enabling outbound email changes the initial intake boundary. It requires a destination rule, sender authority, payload review, retry behavior, and examples where the message must remain unsent.

Do not quietly broaden the scope because a model can do more. A new capability is a product decision that the owner should understand.

A wording adjustment can use a lighter review if it leaves behavior and permissions intact. Even then, replay relevant cases when the instruction change could alter routing or escalation.

Build the smallest agent your client agreed to.

Use Pickaxe to configure the job, knowledge, and actions, then check the agreed examples.

Get started →

A proposed five-day pilot with a concrete deliverable

This is a suggested schedule for the fictional intake assistant, not a promised implementation time. Extend it if the required integrations or policy decisions are more complicated.

  1. Day one: agree on the job, owner, non-goals, and source rules. Produce the initial spec and visible open questions.
  2. Day two: write representative cases, including missing inputs and prohibited actions. Resolve policy questions before connecting write tools.
  3. Day three: build the draft-only agent and validate the package format. Record the configuration and integration behavior.
  4. Day four: replay the cases, inspect outcomes, and fix the failures. Have the client review examples against the spec.
  5. Day five: decide whether to release, keep the pilot bounded, or revise scope. Record the actual decision and remaining limitations.

The deliverable is a package: agreed spec, representative cases, configured agent version, evaluation evidence, and an operator’s pause and recovery instructions. “The chatbot looks good” is too vague for a handoff.

If a critical boundary fails, leave the affected capability disabled while resolving it. A pilot ending with a well-defined blocker can be more useful than a deployment nobody knows how to supervise.

Questions I would settle before the first release

Is a spec just a longer prompt?

Some content overlaps, but the responsibilities differ. The spec records the agreed behavior and evidence requirements; the prompt, source configuration, permissions, and integrations implement them.

Can AI write the spec for me?

It can organize a rough brief and surface ambiguities. The business owner still needs to choose policies and approve boundaries, because fluent wording does not establish permission.

How much detail is enough?

Enough to distinguish an acceptable outcome from an unacceptable one for the important cases. Start with a readable summary and expand where real ambiguity or consequence appears.

Does a passing test suite guarantee reliability?

No. It provides evidence about the covered situations and tested configuration; continue monitoring actual behavior, investigate failures, and add cases for newly discovered conditions.

Who owns updates after handoff?

Name a business owner for behavior decisions and an operator for deployment changes. Record which changes require renewed approval so maintenance does not depend on someone remembering an old chat.

Start with the next misunderstanding you want to prevent

I would begin with one useful job and one situation where the wrong interpretation could matter. Write the expected behavior, identify the implementation control, and define the evidence a reviewer will inspect.

Then build the smallest version that satisfies that agreement. A good spec earns its place by making a disputed decision clear, not by producing more documentation.

If you are building that first client agent in Pickaxe, keep the template beside your configuration and evaluation cases. You will have a clearer starting point for the build and a more useful artifact to hand over when it changes.