All articles
EngineeringOctober 4, 2026

State management in long-running trading systems

Long-running automation depends on knowing what happened before, what is happening now and what should happen next.

One of the first things automated trading taught me is that decisions rarely exist in isolation.

A system cannot simply look at the current situation and forget everything that happened before. It needs context: what it already did, what is currently active and which assumptions remain valid.

That context is state.

And in a long-running trading system, state is not a detail. It is part of the product's memory.

A trading system has memory

An automated action immediately creates new information.

An order may be pending, completed, rejected or waiting for another part of the system to react. The next decision cannot safely ignore that history.

A long-running system therefore behaves less like a collection of independent functions and more like a continuous process with memory.

This sounds obvious until the application restarts.

Restarts do not reset reality

Software eventually restarts, deliberately or unexpectedly.

The external world does not reset with the application. Existing orders, previous decisions, balances and open positions still matter. A robust system needs to recover its understanding of reality rather than assume it is starting from zero.

This is where shortcuts become dangerous.

If the platform does not know what happened before the restart, it may duplicate work, ignore active exposure or make decisions from incomplete information.

Distributed systems make it harder

The challenge grows when responsibilities are split across components.

Different parts of a platform may observe events at slightly different moments. Messages can arrive later than expected. One component can restart while another continues normally.

The objective is not to make everything instantaneous.

The objective is to keep the overall state coherent enough that incomplete information does not lead to unsafe assumptions.

State needs explicit design

State management becomes fragile when it is treated as something that simply happens along the way.

Some information needs to be stored. Some needs to be recomputed. Some needs to be confirmed against an external source before the system can trust it.

Those decisions shape how the platform behaves under pressure.

For Oblivion, this changed how I thought about persistence, recovery and observability. The question was not only what the system should do when everything is normal. It was also what the system should remember when normal conditions disappear.

Looking back

State management is not the kind of feature users normally notice.

Ideally, they never need to think about it. But if a long-running system loses its context, automation quickly becomes guesswork.

Financial software is one of the last places where I want the platform to guess.

State management in long-running trading systems