1. Your stages¶
A stage is the one place in StageFlow where your code runs. Everything else — the branching, the loops, the error handling, the frame — is the graph's business, and the graph is JSON somebody else may well have drawn.
from stageflow import BaseStage, register_stage
@register_stage("LoadTicketStage")
class LoadTicketStage(BaseStage):
"""
description: "Takes a prepared ticket out of data/tickets.json"
icon: "/icons/ticket.svg"
arguments:
ticket_id:
type: string
optional: true
description: "Ticket id (T-1001); empty — the first one in the file"
outputs:
text:
type: string
description: "The customer's message"
customer_id:
type: string
description: "Who wrote it"
"""
category = "support.data"
async def run(self):
wanted = (self.get_arguments().get("ticket_id") or "").strip()
...
self.set_outputs({"text": ticket["text"], "customer_id": ticket["customer_id"]})
The docstring is not documentation that happens to be parsed — it is the
spec, and the editor draws the card from it: the name, the icon, the colour,
which ports exist and what each one is for. Two of the card's fields sit
beside it as class attributes rather than docstring keys — category and
timeout — and Stage specification has the full
grammar with that split. What matters here is that all of it lives in the file
it describes, a few lines from the code it is about.
Every description here is prose, and prose may be written in more than one
language — description: {en: …, ru: …} instead of a string, or a gettext
catalog for a backend with too many stages to repeat themselves. get_specs()
then serves every language at once and the editor picks for its reader, so
nothing about this endpoint or this docstring has to know which language anybody
wanted (Localization).

Small, and about one thing¶
The stages in the example are deliberately tiny: load a ticket, classify it,
search the knowledge base, render a reply, send it, escalate. That is not
neatness. A node on the canvas is a stage, and a stage that both looks
something up and decides what to do with it becomes a node whose meaning has
to be guessed from its name. The decisions belong in the graph, as a
condition or a switch, where they can be seen and argued with.
The test is easy to apply: if the docstring needs the word "and", it is two stages.
Arguments come from the frame¶
A stage never reads the frame. It is handed arguments, and the graph says where they came from:
{
"id": "classify",
"type": "stage",
"stage": "ClassifyByRulesStage",
"arguments": {"vars": {"text": "text", "subject": "subject"}},
"outputs": {"topic": "topic", "urgency": "urgency"},
"next": "search"
}
arguments.vars maps an argument name to a frame variable, arguments.const
to a literal, and a key ending in . makes the value a
CEL expression. outputs maps the other way. The stage
sees a plain dict and returns a plain dict, which is what makes it testable
without a pipeline at all — and what lets the same stage appear twice in one
graph reading different variables.
Errors are roads¶
A stage that raises tells the graph what happened, and the graph decides. That only works if the exceptions are distinguishable:
class LlmAuthError(RuntimeError): # no key — fall back to the rules
class LlmRateLimited(RuntimeError): # 429 — a retry with a pause helps
class LlmUnavailable(RuntimeError): # 5xx — a retry helps too
class LlmBadAnswer(RuntimeError): # the model answered nonsense
Four classes rather than one, because in a graph they become four different
roads: retry on the two transient ones, an except down to the keyword
rules on the first, a different ending for the last. A single bare Exception
would leave the author of the pipeline nothing to branch on — see
Errors and retries.
Events, and text as it is written¶
A stage can emit events, and they reach the editor's log as the run goes:
allowed_events = [
EventSpec("reply_sent", "The reply left for the customer",
payload_schema={"ticket_id": str, "channel": str, "chars": int}),
]
self.emit("reply_sent", {"ticket_id": ticket_id, "channel": "email", "chars": len(reply)})
Declaring them is worth the three lines: an undeclared type is refused, so a typo in an event name is caught rather than silently producing an event nobody listens for.
One payload shape is a contract rather than a convention. Anything with
{"stream": true, "text": "…"} is treated by the editor as a piece of text
the node is writing right now, and shown in a column of the debug panel as it
arrives:
for word in reply.split(" "):
self.emit("reply_chunk", {"stream": True, "text": word + " ", "label": "Reply"})
The editor knows nothing about support bots or language models here — it knows
the payload shape. A transcription stage or a build-log stage gets the same
treatment for free, and label is what the column is called.
Registering them¶
register_stage puts the class in a process-global registry, which is
what makes get_stages() — and therefore /api/stages — possible at all:
Global is also the sentence with consequences mentioned on the index: every stage this process imports is a stage every pipeline in it can name. That is fine while you write the pipelines; it stops being fine the moment somebody else does, which is what step 3 is for.
Next: the endpoints — how the editor reaches any of this.