A decision matrix for teams whose best AI users keep hitting their usage limits by Wednesday.
A working tool. Shabnam Aggarwal Advisory.
Describe the task the way you'd describe it to a colleague. We'll place it on the matrix and tell you which model to use, and why.
The pattern shows up in every organisation we work with. The heaviest users run everything on the biggest model, because it gives the best answer. Then the weekly limit runs out midweek, the extra-usage bill starts climbing, and the instinct is to buy more capacity. Buying more works. Choosing better works first, and costs nothing.
The price gap is real. Anthropic's own API prices are the cleanest public proxy for how fast each model consumes a shared limit: Haiku runs $1/$5 per million tokens in and out, Sonnet $3/$15, Opus $5/$25, and Fable $10/$50. Call it roughly 1x, 3x, 5x, 10x. A recurring weekly workflow moved from Fable to Sonnet frees roughly two thirds of what it was consuming, with no visible quality drop for routine work. Fable also thinks harder by default, so the gap in practice is often wider than the sticker ratio.
Two questions place almost any task. How hard is the thinking, and how long will the model be working? Everything below hangs off those two.
And there are two levers for fixing the problem. Admin controls that set defaults and caps (the stick, covered at the bottom of this page), and helping each person place their own tasks well (the carrot, which is everything in between).
Hard thinking with a clear finish line. Multi-document synthesis, decision memos, one-off analysis where the reasoning quality shows.
The hardest long-running agentic work, a few times a week at most. Deep research, complex builds that run for hours. It burns limits fastest, so make it earn the seat.
Lookups, formatting, short summaries, quick emails. Near-instant, and it barely touches your limit.
The everyday default, and the ceiling for anything that repeats. Drafting, editing, dashboards, meeting notes, monthly reports.
Reach for it when the task has a known shape: look something up, reformat, summarize a short document, draft a routine email, check a list against another list.
Skip it when the task needs weighing or synthesis. You'll spend more time correcting than you saved.
Reach for it when you're doing the actual work of the week: drafting, editing, analysing a spreadsheet, building or refreshing a dashboard, turning meeting notes into actions. Anthropic recommends it as the organisational default, and we agree.
The one hard rule: anything that recurs weekly or monthly runs here or below. More on why in the habits below.
Reach for it when the reasoning quality will be visible in the output: synthesizing many documents into a judgment, a decision memo with real tradeoffs, a hard one-off analysis, a complex build.
Skip it when Sonnet hasn't actually disappointed yet. Escalation is a response to evidence, and it works in both directions.
Reach for it when the task is both judgment-heavy and long-running: deep research across a whole evidence base, an agentic build that runs for hours on a genuinely hard problem. A handful of tasks a week deserve this.
Know the cap: on Premium and Max seats, Fable can use up to 50% of your weekly limit before it starts drawing paid credits. It is deliberately rationed. Treat it that way.
These are drawn from the workflows we see across foundation and nonprofit teams: portfolio leads, program officers, operations, and grants. Find the person who looks most like you and steal their column.
Manages a portfolio of organisations, meets funders, builds her own tools, and is usually the person capping out the team's limits by Wednesday.
| Build a portfolio-wide cost-effectiveness comparison across every org, from source documents up | Fable | Judgment-heavy and long-running: the one build a quarter that earns the top model. Do the heavy sessions here, then maintain it on Sonnet. |
| Program survey data into a funder-relationship base and keep it linked to the portfolio | Opus | Complex one-off build; the monthly upkeep afterward is Sonnet work. |
| Before a funder meeting: who across our portfolio has or wants this funder? | Haiku | A lookup against data that already exists. Seconds, and it barely registers on the limit. |
| Synthesize a grantee's milestone documents ahead of a renewal conversation | Sonnet | Happens every renewal cycle, so it lives at Sonnet. Escalate a single tricky renewal to Opus when the call is genuinely hard. |
Reads proposals, writes decision memos, tracks an evidence base, and answers to an investment committee.
| Weigh a major proposal against the alternatives and draft the decision memo | Opus | The reasoning is the product, and the stakes are visible. Form your own view first, then let the model pressure-test it. |
| Scan the recent literature in your field and flag what changed this quarter | Fable | Deep research across a wide evidence base, quarterly at most. If you find yourself doing it weekly, it moves to Opus. |
| Edit a grant write-up for clarity and length | Sonnet | Bread-and-butter editing. The quality difference above Sonnet is invisible here. |
| Refresh the monthly portfolio dashboard and flag stale updates | Sonnet | Recurring, so Sonnet is the ceiling. The build was the hard part; the refresh is a pattern. |
Runs compliance, files, calendars, and the systems everyone else depends on. Rarely the loudest AI user, often the highest-value one.
| Design the compliance document-collection workflow: what to collect, when to chase, what to check | Opus | Design once with the heavy thinker. The weekly chasing and checking it produces runs on Sonnet or below. |
| Draft the chase emails for missing grantee documents each week | Haiku | Known shape, high frequency. Exactly what the quick model is for. |
| Plan a file-system migration and propose the new structure | Opus | One-off with real tradeoffs. |
| Turn a scheduling mess into a matrix of who can meet when | Sonnet | Structured and slightly fiddly. Sonnet handles it; nothing above it would do visibly better. |
Processes reports, produces the recurring numbers, and holds the line on what the data actually says.
| Summarize incoming grantee reports against stated milestones | Sonnet | Recurring intake work. One caution that has nothing to do with models: read the report yourself before the summary shapes your view. |
| Produce the monthly and quarterly grants reports that are manual today | Sonnet | The definition of a recurring workflow. Build the pattern once, run it cheap forever. |
| Check a batch of records for missing fields before a board cutoff | Haiku | A checklist at volume. Haiku eats this for breakfast. |
| Mine three years of grants data for patterns worth a board conversation | Opus | Genuine analysis with judgment in it, a few times a year. Worth the heavier model when it happens. |
Pick the bigger model because the output disappointed, never because the task feels important. Importance is about stakes; model choice is about difficulty.
The single highest-leverage rule in this whole page. One weekly workflow left on Fable can quietly consume more than a month of everyone else's careful choices.
Long conversations re-read everything above them on every turn, whichever model you're on. Ask related questions together in one message rather than one at a time. Agentic sessions in Cowork burn faster than chat by design: the same task in fewer, tighter sessions costs less than one sprawling one.
If you paste the same background document into chats every week, move it into a shared Project or skill once. You stop paying to re-send it, and your colleagues inherit it.
Design a workflow with Opus or Fable once, then hand the running of it to Sonnet or Haiku. The thinking was the expensive part. Don't keep paying for it after it's done.
The bigger models each have an effort dial, low to max, and it moves usage as much as swapping models does. Save high effort for the task where the reasoning has to hold up, and drop back down for the routine ones, whichever model you picked. Haiku is the exception: no dial, always quick.
Cowork's workflow tool can run a dozen agents in parallel, each burning usage on its own account regardless of the model behind it. Ie. splitting 40 grantee reports across ten sub-agents is fast, but it spends like ten sessions, not one. Reach for it when the deadline is worth the multiplier, not as the default way to work through a list.
A framework only works if people remember it midtask, and most people won't. On Team and Enterprise plans, three admin settings can do the remembering for them, and one measurement keeps the whole thing honest.
Owners can set org-wide instructions that reach every conversation. Paste a condensed version of the matrix and ask Claude to flag, midchat, when a task looks mismatched to the model it's running on.
Org-wide or per role. New conversations start in the right place, and escalating becomes a deliberate choice rather than a habit.
The one hard control of the three: admins can switch off the heavier models and cap the maximum effort level for roles that don't need them, while the roles doing genuinely heavy work keep everything unlocked.
The admin dashboard exports per-user spend by model. The fair metric is each person's share of their own tokens on Sonnet or below, and its month-over-month change. An award for most-improved mix, above a minimum usage floor, lets your heaviest builder win it in the same month they top total spend. Publish the team-level number; keep the individual rows private.
Model guidance for our team When a conversation starts, quietly check whether the model selected fits the task. If there is a clear mismatch, suggest the better fit in one sentence at the top of your first reply, then carry on with the task either way. One suggestion per conversation at most. If the person says "stay on this model," respect that for the rest of the conversation. Default to concise replies. Go longer when the work genuinely needs it. Place the task with two questions: how hard is the thinking (routine vs judgment-heavy), and how long will the work run (a quick exchange vs a long agentic session or a recurring workflow)? Haiku: quick routine tasks. Lookups, formatting, short summaries, routine emails, checking a list against another list. Sonnet: the everyday default, and the ceiling for anything recurring. Drafting, editing, data analysis, dashboards, meeting notes, weekly or monthly reports. If a recurring workflow is running on Opus or Fable, flag it. Opus: judgment-heavy but bounded work. Synthesis across many documents, decision memos, hard one-off analysis. Suggest it only when Sonnet's output has actually disappointed, and suggest dropping back down afterwards. Fable: the hardest long-running agentic work only. Deep research, complex multi-hour builds. It consumes shared usage fastest, so a handful of tasks a week should earn it. Watch two more dials. Effort: the bigger models have an effort setting, and for routine work low or medium stretches usage further. Fan-out: workflows that run parallel sub-agents spend usage per agent, so before fanning out across a list on a heavy model, flag the multiplier and ask whether Sonnet or Haiku agents would do. Two rules override all of this. Our data policy wins regardless of model: [one line on your data rules, ie. no grantee-identifying data outside approved tools]. And judgment stays human: funding decisions, performance conversations, and anything in our own voice to a partner start with a person's thinking, with AI sharpening it afterwards.
One bracketed line needs your own data rule before it goes in, and swap "partner" for your word: grantee, client, student. It fits well inside the 3,000-character limit, so there's room to add more of your own. The same text works in a personal profile's preferences for anyone whose organisation hasn't set it up, and Claude will make the suggestions in-chat either way. One honest caveat: this is a young method. I haven't yet watched it run across a whole organisation for a quarter, so treat it as a working draft, and tell me what you learn.
Model facts and prices from Anthropic's choosing-a-model guide, model overview, Enterprise consumption guide, Fable plan notes, effort settings, and admin model controls, checked August 2026. Model lineups move; treat specifics as current rather than settled. Happy to talk any of this through: aggarwal.shabnam@gmail.com