3.1 · Lesson 1 of 8
Evaluate tool/agent configuration for capability bloat
What You Need to Know
Capability bloat is when an agent's tool and permission surface grows faster than the workflows that justify it. Extra tools feel like progress in design reviews, but at runtime they compete for selection, widen the blast radius of mistakes, and make authorization hard to reason about. Professional architects treat the toolkit as a scarce design surface — not a dumping ground for every integration the org owns.
Bloat shows up as overlapping names ("search" vs "find" vs "lookup"), pilot tools that never graduated, and mega-agents that can both read tickets and mutate billing. The model then spends capacity choosing among near-duplicates instead of solving the user problem. Prompt wording cannot fully compensate for a noisy choice set.
Signals of bloat
- Wrong-tool rate rising while task definitions stay stable
- Tools with zero or near-zero production invocations
- Multiple tools that differ only in naming or thin wrappers over one API
- Read and irreversible write tools sharing one agent identity
Decision rules
- Inventory before you add: every new tool needs a workflow owner and a measured use case.
- Prefer remove or merge over longer instructions when selection accuracy falls.
- Split agents by role when write/admin and read paths share one oversized set.
- Least privilege removes unused capability; logging and confirm dialogs do not substitute for removal.
- Re-eval tool selection after every toolkit change with fixed scenarios.
Why bloat hurts more than it helps
Tool selection is a classification problem over a noisy label set. When labels overlap ("search" vs "find") or include never-used writes, the model spends attention disambiguating instead of solving. Wrong selection is not only an accuracy issue — for irreversible tools it is a safety incident. Architects therefore treat toolkit size as a first-class non-functional requirement with an explicit upper bound per role.
Review checklist
- Is every tool owned by a named workflow?
- Do any two tools differ only by thin wrappers or synonyms?
- Are read and irreversible write tools co-located on one agent?
- What percentage of tools have zero calls in the last 30 days?
- If we removed half the tools tomorrow, which workflows would break?
Exam stems often tempt you to "add guidance" or "confirm every call." Those controls can be useful later, but they do not shrink the capability surface. The professional move is configuration change: remove, merge, or split — then re-measure selection.
Exam application
Prefer answers that shrink or partition the toolkit when mis-selection rises. Distractors include adding tools, lengthening prompts without config change, raising temperature, or equating confirmation dialogs with least privilege. Measure wrong-tool rate before and after.
Exam traps
Treating more tools as more capability
Each overlapping tool increases mis-selection risk. The exam favors shrinking or partitioning toolkits over expanding them when accuracy drops.
Fixing bloat only with longer prompts
Instructions help disambiguation marginally; they do not remove unused destructive or duplicate surfaces. Prefer configuration change first.
Equating logging or confirmation with least privilege
Observability and HITL are complementary controls. Least privilege removes unused capability rather than approving it after the fact.
One mega-agent with every integration
Architects should split by role when write, admin, and read paths share one oversized toolkit.
Practice scenario
A support agent ships with 38 tools: three near-identical ticket search variants, unused CRM write APIs from an abandoned pilot, and a generic "run any SQL" tool. Tool-selection accuracy is falling. What should the architect do first?
Build exercise
Shrink a bloated support-agent toolkit
35 minutes
What you'll learn
- Detect capability bloat from traces and blast-radius tags
- Consolidate duplicates and remove unused tools
- Split high-risk writes onto a narrower agent path
- Prove the change with a tool-selection eval
Step 1
Inventory tools by role, frequency, and blast radius
List every tool on the agent. Tag each as read vs write, how often it is actually invoked in traces, and what damage a wrong call can do.
Why: You cannot justify a toolkit without usage and risk evidence. Exam items reward measuring whether each tool earns its place.
You should see: A table: tool name, owner system, last 30-day call count, read/write, max blast radius.
Step 2
Consolidate look-alike tools into one interface
Replace three search variants with a single search tool that takes a typed mode or filter argument. Delete the unused CRM write APIs.
Why: Near-duplicates are a classic selection-noise pattern. Consolidation reduces the decision space without losing needed capability.
You should see: One search tool definition; deleted unused write tools; agent config diff showing a smaller tool list.
Step 3
Split write/admin tools onto a dedicated agent or gated path
Keep the customer-facing agent on read and draft tools. Move refunds and admin mutations to a separate agent or a gated workflow with stronger authZ.
Why: Capability bloat often mixes high-risk writes with broad reads. Role split is an architectural fix, not only a prompt fix.
You should see: Two agent configs with non-overlapping tool sets and an explicit handoff contract.
Step 4
Re-measure tool selection on a fixed eval set
Run the same scenarios before and after the shrink. Track wrong-tool rate, unused-tool count, and task success.
Why: Architects justify configuration with evidence. The exam expects you to close the loop after a toolkit change.
You should see: Eval table showing wrong-tool rate down and success stable or improved.
Sources
- Tool use overview — docs.anthropic.com
- How to implement tool use — docs.anthropic.com
- Agent Skills overview — docs.anthropic.com — modular capability packaging