Predicts which part of an AI pipeline is about to take the rest down with it. Illustrative simulation, not the trained model.
Pipeline graph
The demo pipeline the package ships with, so there is something to predict against on a laptop.
- 1planner agent
- 2retriever tool
- 3vector store
- 4research agent
- 5model endpoint
- 6synthesizer agent
Cascade risk by node (sample)
The model endpoint is flagged before the nodes downstream of it turn red. Predicting the cascade, rather than tracing it afterwards, is the whole point.
what one observed call carries CallEvent
run_id / scenario / stepstring · string · inta run is a sequence of steps, so risk can be scored per stepcaller / calleestringthe edge the call createscaller_type / callee_typeenumagent · tool · model_endpoint · vector_storelatency_msfloaterror / retriedbooleantoken_costfloatNode and edge features share one order — latency_ms, error_rate, retry_rate, token_cost — so a node's features are just its incoming edges aggregated. One convention, no translation layer.
Node status
| Node | Type | Latency | Errors | Retries | Risk |
|---|---|---|---|---|---|
| planner_agent | agent | sample | 0 | 0 | low |
| retriever_tool | tool | sample | 2 | 1 | watch |
| vector_store | vector_store | sample | 1 | 0 | watch |
| research_agent | agent | sample | 0 | 0 | low |
| primary_model | model_endpoint | sample | 7 | 4 | high |
| fallback_model | model_endpoint | sample | 0 | 0 | low |
| synthesizer_agent | agent | sample | 0 | 0 | low |
Signals come from instrumentation around the pipeline, so a node does not have to cooperate to be watched. An edge with no history yet — the fallback model before a fallback has ever fired — starts from a nominal healthy default rather than a zero.
primary_model
Node
- Name
- primary_model
- Type
- model_endpoint
- Risk
- high
- Fallback
- fallback_model
Predicted impact
- Downstream
- research_agent, synthesizer_agent
- Likely effect
- latency and fallback-quality impact
- Scored at
- step 41 of 60
- Store
- score_history
How a prediction is made
Serving loads a trained GNN and scores a graph snapshot; the alert copy names the node and what it is expected to take with it.
- 1instrument the app
- 2ingest call events
- 3build a graph snapshot
- 4score with the GNN
- 5check input drift
- 6evaluate alert rules
- 7dispatch
This run
- step 41Risk scoredserving
risk_score written to score_history, keyed by run_id and step
- step 41Rule matchedalerting
score at or above threshold · one alert per node, not per call
- step 41Webhook sentdispatch
recorded in alert_history with the score that fired it
- step 44Degradation observedlabeling
written to incident_labels — which is what makes the prediction checkable
Fault injection against the demo pipeline is what generates the labelled failures the model trains on. Without labels there is nothing to train and nothing to grade.
Input-distribution drift
A customer's topology changes as they ship new agents and tools, so the model's own inputs drift. Worth checking before that becomes a reliability problem for the predictor itself.
PSI by feature (sample)
How to read it
- under 0.1no meaningful drift
- 0.1 to 0.2moderate
Worth watching; the model is still in the distribution it was trained on.
- over 0.2significant
The served model is being asked about a pipeline that no longer looks like its training set. Retrain.
Population Stability Index against quantile bins fixed at training time, computed over numpy. Serving never needs the raw training data, only the reference distribution stored next to the model.
Fault injection & training
You cannot learn a cascade from a healthy pipeline. The scenarios below inject the failures that produce the labels.
| Scenario | What it injects | Cascade expected |
|---|---|---|
| baseline | nothing | no |
| rate_limit_model | model endpoint throttling | yes |
| vector_db_degradation | retrieval quality decay | yes |
| cost_spike_model | token cost blow-out | yes |
| vector_store_flaky | intermittent store errors | yes |
| compound_cascade | two faults at once | yes |
Training run
- 1run scenarios
- 2record call events
- 3label affected nodes
- 4build graph dataset
- 5train the GNN
- 6persist model + drift reference
A fault ramps in over several steps rather than switching on, because a real rate limit does not arrive as a step function. A baseline model is kept alongside the GNN so the graph structure has to earn its place.
Alert rules & history
| Node | Type | Score | Channel | Fired |
|---|---|---|---|---|
| primary_model | model_endpoint | sample | webhook | sent |
| retriever_tool | tool | sample | webhook | below threshold |
| vector_store | vector_store | sample | webhook | below threshold |
What the alert says
- model endpointimpact copy
expect latency and fallback-quality impact
- vector storeimpact copy
expect downstream generation quality to degrade soon
- agentimpact copy
its downstream tool and model calls may start failing
- toolimpact copy
downstream steps depending on it may start failing
Alerting is off until a webhook is configured, and Slack and PagerDuty are both webhook-shaped, so one delivery path covers all three. Every send is written to alert_history with the score that caused it.
Pointing it at a real pipeline
- 1cascaid run -- python your_app.py
- 2stack auto-detected
- 3events written per run
- 4cascaid ingest --follow
- 5risk in the dashboard
| Layer | Detected | Adapter |
|---|---|---|
| Orchestrator | LangGraph | supported |
| Orchestrator | CrewAI | supported |
| Model gateway | LiteLLM | supported |
| Vector store | Pinecone · Weaviate · pgvector | supported |
Detection reports what it finds rather than picking a winner, so an orchestrator that is not LangGraph is not permanently second-class. Instrumentation needs no changes to the app itself.
Also exposed
- MCP serverget_cascade_risk
Any agent can ask what the current cascade risk is for a run, rather than a human reading a dashboard.
- Grafana data sourcesearch · query
For teams that already have a wall of dashboards.
- FastAPI endpointstoken auth
/risk/{run_id}, /pipeline/{run_id}, /track-record/{run_id}.
What gets persisted
Postgres in production, SQLite for tests and the demo, one schema either way.
data model score_history · incident_labels · alert_history
run_id / step / node_namestring · int · stringindexed, because every question is asked per run and per noderisk_scorefloatwhat was predicted, when it was predictedincident_type / occurred_at / sourcestring · timestamp · stringwhat actually happenedmessage / channel / sent_atstringwhat was sent, so an alert can be auditedconfigkey / valuethreshold and webhook live here, not in a releaseauth_sessiontoken / expires_atthe dashboard's own sessionsPredictions and incidents are stored separately on purpose. Keeping them apart is what lets the track record be computed rather than asserted.
Track record
Whether the predictions held up, which is the only claim worth making about a predictive tool.
Sample prediction accuracy over runs
Run it yourself
- pipx install cascaidinstall
Or uv tool install cascaid.
- cascaid demozero setup
Spins up the synthetic fault-injection pipeline, trains on it, and seeds a local SQLite store. No Postgres, no Docker, nothing to configure first.
- docker compose upthe real thing
Dashboard, model server and Postgres in one command.
Published on PyPI. The walkthrough above is a wireframe of the dashboard, not the trained model.