NeuroRoute NeuroRoute SISL CloudWorx
Login
03Capabilities/ 9 chapters

Strategies, coverage, and control

Routing is a policy decision, so the policy surface is the product. Six strategies, ten providers, and per-key controls that let platform teams delegate access without losing the guardrails.

At a glance
Routing strategies6
Providers connected10+
Models under management38
Health check interval15 s
03.1

Six strategies, set globally or per key

A strategy is a weighting over the same four scores. Set one default for the organisation and override it per API key, per route, or per tenant.

Task-Aware

The default. Classify, then apply the weights tuned for that task type. Best all-round savings without a quality cliff.

Cost-First

Cheapest model that satisfies the hard constraints. For high-volume, low-stakes traffic where budget is the binding limit.

Quality-First

Highest benchmark score within a cost ceiling you set. For customer-facing output where a bad answer costs more than tokens.

Latency-First

Lowest trailing p95 among qualified models. For interactive surfaces with a response-time budget.

Cascade

Run the ladder cheapest-first, escalate on a failed quality check. Deepest savings on mixed-difficulty workloads.

Manual Pin

Force a specific model for a key, route or tenant. For regulated workloads and A/B baselines.

03.2

Model and provider coverage

New models are benchmarked and priced before they are eligible for routing. Providers in a degraded state are excluded from candidate sets until they recover.

ProviderModels in rotationRegionsStatus
OpenAIGPT-4o, GPT-4o mini, o3-miniUS, EUOperational
AnthropicClaude Sonnet 4, Claude Haiku 3.5US, EUOperational
GoogleGemini 2.0 Flash, Gemini 1.5 ProUS, EU, APACOperational
GroqLlama 3.3 70B, Llama 3.1 8BUSOperational
MistralMistral Large, Mistral SmallEUOperational
xAIGrok 3USDegraded
Failover is part of routing, not a separate feature. When a provider's error rate or latency crosses your threshold, its models leave the candidate set and traffic re-routes on the next request — no retry storm, no code change.
03.3

Observability and control

Everything the router decides is queryable, and everything it is allowed to decide is configurable.

Decision-level tracing

Look up any request by ID and get the classification, the full candidate table, the cascade path, and the cost proof.

Spend controls

Monthly caps and rate limits per API key, with soft-warning thresholds and automatic downgrade instead of hard failure.

Model preferences

Allow-lists, per-task preferred and fallback models, and quality floors expressed as a single number your team can reason about.

Export and integration

CSV and JSON export, per-request webhooks, and an OpenTelemetry span for every routing decision.

Routing Rules, in the admin dashboard: per-task-type routing, with explicit cascade chains where you want them and score-based Auto everywhere else. Reproduced here as live markup rather than a screenshot, so it stays legible at any zoom and in either theme.