Claude Cowork + Project · Claude Cowork + raia Agent · Claude Opus 5 pricing

Stuff the context, or fetch what matters

Every prompt in Claude Cowork against a Project re-reads every uploaded page, for every user. Connect the same Claude to a raia agent over MCP and the agent searches one shared vector store, distills the matches with a smaller model, and hands back only what the question needs. Drag the sliders to see what that does to tokens, cost, and headroom.

100 documents1,000 pages · 500K tokens
25 userssame knowledge base
401,000 prompts a month in total

Claude Cowork + Project

Documents live in project knowledge. Each user's prompts carry the whole library into context.

Full context
Idle ⚠ exceeds project capacity
Input tokens / promptlibrary + prompt
Output tokensthe answer
Cost / promptOpus 5 rates
Project capacity used
Context window is the ceiling. Past it, the project cannot hold the library.

Claude Cowork + raia Agent

Documents live in one raia knowledge base. Claude calls the agent over MCP; it searches, distills, and returns only what is needed.

raia MCP
Idle
Input tokens / prompt2 calls · distilled reply
Output tokenstool call + the answer
Cost / promptincl. raia's smaller model
raia knowledge base used
One index serves every connected user. Per-prompt tokens do not grow with the library.

Running total

Every "Run one prompt" adds that prompt's tokens and cost to each side. Run it a few times and watch the gap open.

Claude Cowork + Project0 prompts run
$0
total cost so far+$0 next prompt
Input tokens0
Output tokens0
share of the two totals
Claude Cowork + raia Agent0 prompts run
$0
total cost so far+$0 next prompt
Input tokens0
Output tokens0
share of the two totals
Tokens per promptfewer input tokens through raia
Projectraia
Cost per promptsaved on every question
Projectraia
Monthly cost, all userssaved per month
Projectraia
Signal in what Claude readsshare of library-derived input that answers the question
Project
raia
Running it for 25 users
Claude Cowork + Project
Claude Cowork + raia Agent
Copies of the library to keep current
25 project uploads to refresh when a document changes
1 knowledge base in raia Command; every connected user sees the update
Spending limits
No per-user cap on how much context each prompt consumes
Hard monthly token caps per agent, usage dashboard, stops at the limit until an admin raises it
Access and control
Each user manages their own project and connectors
One OAuth connector per user, role-based access, revoke a key and access ends immediately
Other systems (CRM, tickets, drive)
Exported and re-uploaded by hand into each project
30+ connectors sync into the same agent; the MCP call reaches all of them

Cost per prompt as the library grows

Project cost climbs with every page. raia stays flat because the knowledge base, not Claude's context, holds the library, so a prompt reads the same few passages whether the index holds 10 documents or 2,000. Switch to a log scale to see both lines at once.

Claude Cowork + Project beyond project capacity Claude Cowork + raia Agent
Assumptions edit any value, everything recomputes
Library
Per prompt
raia agent over MCP
Claude Opus 5 pricing (USD per million tokens)

How to read this. Pricing uses Claude Opus 5 API rates: $5 per million input tokens, $25 per million output, cache reads at 10% of input. raia's distilling step is priced at the selected model's rates (Claude Haiku 4.5 by default, $1 in and $5 out) plus raia's markup on those tokens (50% by default), and is included in the raia cost. raia lets you choose the agent's model, so the picker also carries OpenAI GPT-5.6 Luna ($0.20 in, $1.20 out) and Google Gemini 3.8 Flash ($0.75 / $3.75, list price through 2026) and Gemini 3.5 Flash-Lite ($0.30 / $2.50) at their September 2026 API list prices. The open-weight preset is a placeholder rate; replace it with your hosting provider's price. A Claude project or Cowork seat is billed by subscription, so the project cost shown is what the same token traffic would cost at API rates, which is what the platform must absorb behind the seat.

What is left out. One-time embedding and hosting for the raia knowledge base, the raia subscription itself, cache-write premiums on the first prompt, and thinking tokens on either side. Retrieval quality depends on chunking and embeddings; the signal share is an illustration, not a benchmark. Project capacity defaults to the 1M-token Opus 5 context window. Claude.ai and Cowork may enforce a lower project-knowledge limit; lower the value under Assumptions to match your plan.