PURPOSE-BUILT FOR ANTHROPIC CLAUDE

Semantic context routing & memory for Claude Agents.

Stop choking Claude’s 200k & 1M token context windows with noisy dumps. ContextRail AI is pure developer infrastructure that dynamically prunes irrelevant context, aligns prompt caches, and injects persistent cross-session episodic memory in under 50ms.

Connect via Claude MCP
78% Token Reduction
< 50ms Routing Overhead
99.4% Needle Recall
Zero Model Training
LIVE ROUTING ENGINE TELEMETRY
01. Raw Ingestion Window 142,800 tokens
Multi-doc RAG dumps, tool outputs & multi-turn history
↓
02. ContextRail Semantic Pruner -78.2% COMPRESSION
Salience reranker + graph episodic memory retrieval (38ms)
↓
03. Claude Sonnet / Haiku Dispatch 31,120 tokens
Cache-aligned prompt format · 99.4% Needle Recall · TTFT 142ms
ENGINEERING ARCHITECTURE

A high-velocity context highway for autonomous agents

Standard RAG dumps brute-force entire document chunks into the context window, inflating latency and Claude API costs. ContextRail AI introduces a semantic routing layer that filters, summarizes, and routes high-salience tokens.

⚡

Sub-50ms RAG Compression

Prunes syntactic fluff, redundant markdown, and low-salience paragraphs while preserving crucial technical needles. Compresses token footprints by up to 78% with zero hallucination penalty.

🧠

Cross-Session Agent Memory

Equip Claude with persistent episodic memory. Our hybrid vector-and-knowledge-graph index remembers previous user decisions, codebase invariants, and tool results across sessions.

🎯

Prompt Cache Boundary Alignment

Automatically formats prompts so static system guidelines and immutable tool schemas sit cleanly behind Anthropic's 5-minute Prompt Caching breakpoints, saving up to 90% on input token costs.

🛡️

Zero Foundation Model Training

Your enterprise codebase and agent interactions are never used for training. Context is processed in volatile memory partitions with AES-256 tenant isolation and GDPR sovereignty.

🔌

Native Model Context Protocol (MCP)

First-class compliance with Anthropic's open MCP standard. Drop our official server card into Claude Desktop or Claude Code in seconds to expose memory and routing tools immediately.

📊

Token Savings & Latency Telemetry

Inspect per-turn token economy, cache hit ratios, and millisecond routing latency directly from our developer console or via automated webhook alerts.

NATIVE MCP INTEGRATION

Connect to Claude Code & Claude Desktop in 60 seconds

ContextRail AI provides a verified Model Context Protocol server exposing real-time context routing and episodic memory tools directly to your local or cloud Claude runtimes.

OPEN STANDARD · MCP v1

Expose ContextRail directly in Claude

By adding ContextRail AI to your claude_desktop_config.json or Claude Code agent configuration, Claude gains automatic access to semantic memory queries and dynamic context pruning tools.

route_context — Salience reranking & pruning
compress_rag_context — Sub-50ms token reducer
retrieve_agent_memory — Cross-session vector lookup
store_agent_memory — Persistent episodic entity graph
Terminal — one command
$ claude mcp add --transport http contextrail https://api.contextrail.cloud/mcp

$ curl https://api.contextrail.cloud/.well-known/mcp/server-card.json
claude_desktop_config.json
{
  "mcpServers": {
    "contextrail": {
      "command": "npx",
      "args": ["-y", "@contextrail/mcp-server"],
      "env": {
        "CONTEXTRAIL_API_KEY": "cr_live_your_api_key_here",
        "CONTEXTRAIL_ENDPOINT": "https://api.contextrail.cloud/mcp",
        "TARGET_MODEL": "claude-sonnet"
      }
    }
  }
}
EMPIRICAL BENCHMARKS

Measured performance on Claude Sonnet

Evaluated on real-world multi-repo refactoring and 100k+ token financial multi-document Q&A suites. ContextRail AI outperforms vanilla RAG and raw ingestion across all core dimensions.

Routing Strategy Context Reduction Needle Recall TTFT Latency Cost / 10k Agent Turns
ContextRail AI (MCP Engine)
Semantic Reroute + Cache Breakpoint
78.2% compressed 99.4% accurate 142 ms $4.20 USD
Standard Vector RAG (LangChain / LlamaIndex)
Top-K chunk concatenation
32.0% compressed 76.5% accurate 680 ms $18.50 USD
Raw Uncompressed Ingestion
Dumping entire repo/docs into 200k window
0% (Full dump) 84.1% accurate 1,850 ms $38.90 USD
REST API · OPENAPI 3.1

Public HTTP API for programmatic context routing

Every MCP tool is also a documented REST endpoint, so you can call ContextRail AI from any runtime, CI job or backend service. TLS 1.3, bearer-key auth, RFC 9457 error bodies.

Base URL https://api.contextrail.cloud/api/v1 OpenAPI spec & MCP server card →
MethodPathDescriptionAuthRate limit
POST /api/v1/context/route Rank and prune a context payload before injection Bearer key 120 req/min
POST /api/v1/context/compress Sub-50ms semantic RAG chunk compression Bearer key 120 req/min
POST /api/v1/memory/query Cross-session episodic memory lookup Bearer key 60 req/min
POST /api/v1/memory/write Persist agent decisions into the vector-graph index Bearer key 60 req/min
GET /api/v1/metrics/efficiency Token savings, cache hit ratio, routing latency Bearer key 120 req/min
POST /api/v1/webhooks Register an HTTPS sink for budget and latency alerts Bearer key 10 req/min
JSON in, JSON out with Idempotency-Key support on all POST routes. MCP transport: http streamable at https://api.contextrail.cloud/mcp. Scale plan raises limits to 1,000 req/min and per-tenant concurrency of 64.
TRANSPARENT USD PRICING

Predictable SaaS plans for builders and enterprises

No convoluted credit tokens. Transparent monthly billing in USD with instant API key issuance and zero commitment.

Sandbox

For individual developers prototyping agents with Claude Code or local desktop scripts.

$0/mo
  • 10,000 routed tokens/month
  • 1 concurrent MCP agent connection
  • Standard semantic compression
  • Community Discord & GitHub support

Scale

For heavy multi-agent fleets, high-frequency autonomous workflows, and enterprise scale.

$149/mo
  • 15,000,000 routed tokens/month
  • Unlimited concurrent MCP agent fleets
  • Hybrid vector + knowledge graph memory
  • Dedicated priority routing & webhooks
  • 99.9% Uptime SLA & Slack engineering channel
FREQUENTLY ASKED QUESTIONS

Technical specifications & FAQs

Everything you need to know about ContextRail AI integration, security guarantees, and compatibility with Claude models.

How does ContextRail AI integrate with Claude Sonnet and Haiku? +
ContextRail AI operates either as an intelligent API proxy or via our official Model Context Protocol (MCP) server. When your agent prepares to execute a prompt, ContextRail inspects the uncompressed payload, performs salience reranking and cross-session memory retrieval, and formats the output into cache-aligned prompt blocks before delivering it to Claude.
What is the latency overhead of semantic context routing? +
Our routing pipelines execute on European edge infrastructure and optimized GPU inference endpoints with an average overhead of just 38ms to 48ms. Because token volume is cut by ~78%, the net Time-To-First-Token (TTFT) for Claude is actually substantially faster overall compared to ingesting raw uncompressed documents.
Do you train foundation models on our prompts or codebase data? +
Absolutely not. ContextRail AI has an ironclad zero-training architecture. We never train, fine-tune, or calibrate any public or proprietary machine learning models on customer payloads, episodic memories, or prompt queries. Your data is isolated under AES-256 tenant keys and processed in ephemeral memory.
How does ContextRail work with Anthropic's Prompt Caching? +
Anthropic offers up to a 90% discount on cached prompt tokens when exact prefixes are maintained across requests. ContextRail automatically structures your agent prompts so that static system instructions, immutable tool definitions, and long-term memory blocks remain stable at the head of the prompt, ensuring maximum cache hit rates across consecutive agent turns.
Can I use ContextRail with Claude Code and Claude Desktop? +
Yes! ContextRail publishes an official MCP server package (@contextrail/mcp-server) and exposes a verified server card at /.well-known/mcp/server-card.json. You simply paste the snippet into your Claude configuration file and your assistant immediately acquires our routing and memory tools.
Where is ContextRail AI headquartered and where is data processed? +
ContextRail AI is headquartered in Madrid, Spain (Paseo de la Castellana 95). All European user data is processed strictly within European Union data centers (Madrid and Frankfurt regions) in full compliance with the EU General Data Protection Regulation (GDPR) and the European AI Act.
ENGINEERED IN MADRID, SPAIN

Building high-throughput infrastructure for the age of autonomous agents

ContextRail AI Software S.L. is a pure B2B software product engineered and operated by our systems team at Paseo de la Castellana 95, 28046 Madrid, Spain. We believe the biggest bottleneck facing autonomous agents is not model intelligence, but context bandwidth and retrieval bloat.

By treating context as a high-speed routing fabric rather than a static text bucket, we empower engineering teams to run complex multi-agent architectures on Claude with predictable costs, instant retrieval, and cross-session persistence.

Legal Entity ContextRail AI Software S.L.
Tax ID (NIF) B-88514920 (Madrid, Spain)
Headquarters Paseo de la Castellana 95, Planta 18, 28046 Madrid
Primary Stack Anthropic Claude API (Sonnet and Haiku)
Protocol Standard Model Context Protocol (MCP)
Regulatory Alignment EU GDPR & EU AI Act (Zero Retention)
Platform Model 100% Pure Software SaaS

Gregorio Jiménez Salgado

Founder & Lead Memory Architect
Architect of the ContextRail memory fabric. Specializes in episodic graph indexing, multi-turn prompt caching boundaries, and streaming middleware for Claude Sonnet 4.6 and Haiku pipelines.

Marcos Valdés Chen

Co-Founder & Principal Systems Engineer
Expert in high-concurrency memory protocols and MCP tool interfaces. Leads the development of low-latency semantic token compressors and vector cache replication across European regions.