WarEngine: Battlefield State & Superstep Execution
Since: v0.10.0 (Phase 22)
Crates: paladin-battalion (WarEngine, WarGraph), paladin-core (Battlefield,
Waypoint), paladin-storage (the WaypointPort backends)
Every code example targets the current v0.10.0 workspace. The substantive examples are real, compiled code pulled from the
paladin-doc-examplescrate via mdBook{{#include}}, so they are checked against the live API; a few illustrative fragments are markedrust,ignore. The API forms are verified againstcrates/paladin-battalion/andcrates/paladin-core/.
The WarEngine is the superstep engine every Paladin Battalion pattern in this book ultimately
runs on: it executes a WarGraph of nodes over a typed Battlefield, in bounded supersteps
that checkpoint automatically after each one and resume with zero re-execution after a crash.
Unlike the legacy Campaign graph, a WarGraph permits cycles — including self-loops — so
iterative workflows (retry-and-refine, evaluate-optimize loops) are expressible directly. This
page is the substrate the Control Flow, Parley & Chronicle,
Aegis and Agent Runtime guides build on — those guides
cover routing, pause/resume, fault tolerance and middleware respectively; this page does not
re-explain any of them.
Table of Contents
- Building a Graph: Battlefield State and Superstep Merge Semantics
- Waypoint Checkpointing and Addressing
- The Three WaypointPort Backends
- EngineConfig, EngineLimits and Bounded Iteration
- WaypointRetentionService
- The Graph Fingerprint
- Where to Go Next
Building a Graph: Battlefield State and Superstep Merge Semantics
A Battlefield is the typed shared state a WarGraph's nodes read and write. Its shape is
declared once, as a BattlefieldSchema of FieldSpecs — each field names a DispatchRule
(LastWrite, Append, MergeObject, Sum, or Custom) that decides how two nodes' concurrent
writes to the same field within one superstep are merged. A node's contribution is a StateDelta:
a set of field values, never a direct mutation — the engine merges every delta produced in a
superstep into the Battlefield in one step, through each field's own dispatch rule, so the merge
order is deterministic regardless of how many nodes ran concurrently.
The graph below is deliberately small and cyclic: one node self-loops over a (count, status)
Battlefield a few times before falling out of the loop — the same shape that makes cyclic
execution useful for retry-and-refine style workflows. build_graph takes the EngineLimits
to construct the graph with — pass EngineLimits::default() for the built-in bounds, or the
limits configure_limits derives from EngineConfig (see
EngineConfig, EngineLimits and Bounded Iteration
below).
use paladin_battalion::engine::graph::{EdgeSpec, EngineLimits, NodeSpec, WarGraph};
use paladin_battalion::engine::node::{NodeContext, StateNode, StateNodeError};
use paladin_core::platform::container::battalion::campaign::EdgeCondition;
use paladin_core::platform::container::battlefield::{
Battlefield, BattlefieldSchema, DispatchRule, FieldName, FieldSpec, StateDelta,
};
use paladin_core::platform::container::directive::{Directive, NextStep};
use paladin_core::platform::container::waypoint::NodeId;
/// A pure `StateNode` that increments a `count` field each visit and
/// self-loops -- via `NextStep::Edges` and a `Contains("looping")` condition
/// on its own outgoing edge -- until `count` reaches `target`, then writes
/// `status = "done"` and falls out of the loop. The smallest shape that
/// demonstrates cyclic superstep execution, not a straight-line DAG.
struct LoopUntil {
target: u64,
}
#[async_trait]
impl StateNode for LoopUntil {
async fn run(
&self,
state: &Battlefield,
_ctx: &NodeContext,
) -> Result<Directive, StateNodeError> {
let count_field = FieldName::new("count").map_err(|e| StateNodeError(e.to_string()))?;
let status_field = FieldName::new("status").map_err(|e| StateNodeError(e.to_string()))?;
let current: u64 = state
.get(&count_field)
.map_err(|e| StateNodeError(e.to_string()))?
.unwrap_or(0);
let next = current + 1;
let mut delta = StateDelta::new();
delta
.set(count_field, next)
.map_err(|e| StateNodeError(e.to_string()))?;
delta
.set(
status_field,
if next >= self.target {
"done"
} else {
"looping"
},
)
.map_err(|e| StateNodeError(e.to_string()))?;
Ok(Directive {
delta,
next: NextStep::Edges,
})
}
}
/// Build a small cyclic `WarGraph`: one `Function` node self-loops over a
/// `(count, status)` `Battlefield` a few times, then falls out of the loop
/// once `status` reads `"done"` -- a shape `WarGraph::validate` accepts
/// precisely because cycles, including self-loops, are legal (ENG-FR-02),
/// unlike the legacy Campaign graph's cycle-rejecting validation. Takes the
/// `EngineLimits` to construct the graph with -- see [`configure_limits`]
/// for how a deployment derives them from `EngineConfig`, or pass
/// `EngineLimits::default()` to accept the built-in bounds.
pub fn build_graph(limits: EngineLimits) -> Result<WarGraph, Box<dyn std::error::Error>> {
let count = FieldName::new("count")?;
let status = FieldName::new("status")?;
let schema = BattlefieldSchema::new(vec![
FieldSpec::new(
count,
DispatchRule::LastWrite,
Some(serde_json::json!(0)),
false,
),
FieldSpec::new(status, DispatchRule::LastWrite, None, false),
]);
let mut graph = WarGraph::new(schema, limits);
let looper = NodeId::new("looper");
graph.add_node(
looper.clone(),
NodeSpec::Function(Arc::new(LoopUntil { target: 3 })),
);
graph.add_edge(EdgeSpec {
from: looper.clone(),
to: looper.clone(),
condition: Some(EdgeCondition::Contains("looping".to_string())),
});
graph.add_entry(looper);
Ok(graph)
}
Waypoint Checkpointing and Addressing
Exactly one Waypoint is persisted automatically after every superstep — a full snapshot of
the Battlefield as of that point, never an incremental diff. A Waypoint is addressed by the pair
(ThreadId, WaypointId): ThreadId identifies the run, WaypointId identifies one checkpoint
within it, and each Waypoint also carries parent_waypoint_id lineage back to the start of the
thread. A Waypoint's vanguard field — a Vec<NodeId>, not a struct of its own — lists the nodes
ready to execute in the next superstep; an empty vanguard after a superstep is what RunOutcome:: Completed means.
use paladin_battalion::engine::{RunOutcome, WarEngine};
use paladin_core::platform::container::waypoint::ThreadId;
use paladin_storage::waypoint::in_memory::InMemoryWaypointStore;
/// Build the graph, run it to completion over a fresh
/// `InMemoryWaypointStore`, and return the outcome plus the store and
/// thread so [`inspect_waypoints`] can read back the Waypoint history the
/// run left behind -- a `RunOutcome::Failed` carrying
/// `EngineError::RecursionLimitExceeded` is what a graph that never falls
/// out of its loop would produce once `EngineLimits::max_supersteps` is
/// exhausted.
pub async fn run_engine()
-> Result<(RunOutcome, Arc<InMemoryWaypointStore>, ThreadId), Box<dyn std::error::Error>> {
let (limits, durability) = configure_limits()?;
let graph = build_graph(limits)?;
let store = Arc::new(InMemoryWaypointStore::new());
let engine = WarEngine::new(mock_paladin_port(), store.clone()).with_durability(durability);
let thread = ThreadId::new("superstep-engine-guide")?;
let outcome = engine
.start(&graph, thread.clone(), StateDelta::new())
.await?;
Ok((outcome, store, thread))
}
use paladin_core::platform::container::waypoint::WaypointId;
use paladin_ports::output::waypoint_port::WaypointPort;
/// Given the store and thread [`run_engine`] just used, read back the
/// persisted Waypoint via `WaypointPort::latest`, addressed by `(ThreadId,
/// WaypointId)`, and return its `vanguard` -- the nodes ready for the next
/// superstep the engine checkpointed after the run's final superstep.
pub async fn inspect_waypoints(
store: &InMemoryWaypointStore,
thread: &ThreadId,
) -> Result<Option<(ThreadId, WaypointId, Vec<NodeId>)>, Box<dyn std::error::Error>> {
let waypoint = store.latest(thread).await?;
Ok(waypoint.map(|wp| (wp.thread_id, wp.waypoint_id, wp.vanguard)))
}
The Three WaypointPort Backends
Every Waypoint write goes through the WaypointPort trait (paladin-ports), and three backends
implement it, all passing the same shared contract test suite:
| Backend | Path |
|---|---|
| In-memory | crates/paladin-storage/src/waypoint/in_memory.rs (InMemoryWaypointStore) |
| SQLite | crates/paladin-storage/src/waypoint/sqlite.rs |
| Postgres | crates/paladin-storage/src/waypoint/postgres.rs |
InMemoryWaypointStore needs no Cargo.toml feature flag beyond what crates/doc-examples
already declares — paladin-storage's waypoint module is not feature-gated, unlike its
sqlite/mysql/postgres submodules.
EngineConfig, EngineLimits and Bounded Iteration
Because a WarGraph permits cycles, every run needs bounds so it always terminates.
EngineLimits (paladin-battalion) is what a WarGraph is constructed with; the app-facing
EngineConfig (src/config/engine.rs) is what a deployment actually configures, and converts
into EngineLimits via impl From<EngineConfig> for EngineLimits:
use paladin::config::engine::EngineConfig;
use paladin_battalion::engine::WaypointDurability;
/// Configure the engine's bounded-iteration limits and Waypoint durability
/// through the app-facing `EngineConfig` -- the same struct the
/// `APP_ENGINE_MAX_SUPERSTEPS`, `APP_ENGINE_MAX_NODE_VISITS`,
/// `APP_ENGINE_RUN_TIMEOUT_SECS`, `APP_ENGINE_WAYPOINT_DURABILITY` and
/// `APP_ENGINE_MAX_MUSTER_TASKS` environment overrides populate at boot --
/// then convert it into the `EngineLimits` a `WarGraph` is constructed with
/// (`waypoint_durability` stays on the source `EngineConfig` value itself;
/// it is not part of `EngineLimits` and is passed to
/// `WarEngine::with_durability` separately).
pub fn configure_limits() -> Result<(EngineLimits, WaypointDurability), Box<dyn std::error::Error>>
{
let config = EngineConfig {
max_supersteps: 20,
max_node_visits: 10,
run_timeout_secs: Some(60),
waypoint_durability: WaypointDurability::Strict,
max_muster_tasks: 50,
..EngineConfig::default()
};
config.validate()?;
let durability = config.waypoint_durability;
let limits: EngineLimits = config.into();
Ok((limits, durability))
}
EngineConfig field | Bounds | APP_ENGINE_* override |
|---|---|---|
max_supersteps | superstep count before EngineError::RecursionLimitExceeded | APP_ENGINE_MAX_SUPERSTEPS |
max_node_visits | per-node execution count before EngineError::NodeVisitLimitExceeded | APP_ENGINE_MAX_NODE_VISITS |
run_timeout_secs | whole-run wall-clock budget before EngineError::RunTimeoutExceeded | APP_ENGINE_RUN_TIMEOUT_SECS |
waypoint_durability | Strict (a save failure fails the run) or BestEffort (logged, run continues) | APP_ENGINE_WAYPOINT_DURABILITY |
max_muster_tasks | tasks one NextStep::Muster directive may request before EngineError::MusterTaskLimitExceeded | APP_ENGINE_MAX_MUSTER_TASKS |
A graph that never falls out of its own loop hits EngineError::RecursionLimitExceeded once
max_supersteps is exhausted — the same limit EngineLimits::default() sets to 50 and this
page's own example graph would hit if its LoopUntil node never wrote status = "done".
WaypointRetentionService
Waypoint history grows without bound unless something prunes it. WaypointRetentionService
(src/application/services/waypoint_retention.rs) is the application-layer policy: it defines
the single project-wide rule for what may never be deleted — a thread's latest Waypoint, plus
every Waypoint whose status is AwaitingInput — and drives the storage-layer prune() free
function (crates/paladin-storage/src/waypoint/retention.rs) with that rule and a configured
age/count bound. The storage layer itself carries no opinion about what "protected" means; it is
handed the answer as a plain function argument.
The Graph Fingerprint
A WarGraph's content fingerprint — GRAPH_FINGERPRINT_VERSION (currently "v6") plus a
blake3 hash over node ids, edge specs and schema field names — is compared on resume, so
resuming a thread against a structurally different graph fails fast with GraphMismatch rather
than silently replaying stale routing. The fingerprint is deliberately not computed over
prompts or models (those may be hot-swapped without changing run semantics), and it excludes
every EngineLimits field — raising max_supersteps to let a resumed run continue is a
legitimate operator action, not a graph change. The version tag has bumped twice since Phase 22:
v1 → v2 fixed a delimiter-collision encoding bug, and the current v6 covers the Aegis and
structured-output-schema sections the fingerprint's canonical byte stream now includes.
Where to Go Next
This page covers the engine's own state, checkpointing and bounds — not what runs on top of it:
- Routing a node's output to the next node, including dynamic jumps and Muster fan-out — see Control Flow: Dynamic Routing & Subgraphs.
- Pausing and resuming a run, including human-in-the-loop Parleys and graceful shutdown — see Parley & Chronicle.
- Per-node fault tolerance — retry, timeout, error handlers, model fallback and caching — see Aegis: Retry, Timeout, Error Handlers, Model Fallback and Node Caching.
- Middleware, context management, Vault memory and structured output — see Agent Runtime.