Eleven scenarios · eight mechanics · one fictional system

Observe the system.
Decide under constraints.

Northbridge is an imaginary platform that goes through failures, migrations, deployments and load peaks. The laboratory does not test definitions: it asks you to interpret signals, recognise dependencies and choose technical trade-offs.

No leaderboard and no profiling. State remains only in page memory and disappears when you close or reload it.

Infrastructure Lab

From topology to terminal, through to final architecture.

Each scenario uses different information and interactions: diagram selection, ordering, signal budgets, sliders, policy matrices, a simulated terminal and component composition under constraints.

Northbridge Control Room0/11 completed

Method

Useful interaction, not decoration.

Animations represent a state or consequence. Evaluations do not label the person: they describe the solidity of the choice, visibility of the system, residual risk and introduced complexity.

01

Progressive information

Each scenario provides only relevant data and retains some irrelevant signals, as happens during a real incident.

02

Explicit trade-offs

A solution may restore service while leaving risk, or be reliable but unnecessarily complex.

03

Accessibility and privacy

Native controls, keyboard operation, reduced motion and no persistence or transmission of answers.

What to take away

A good technical decision is understandable, controlled and verifiable.

The laboratory does not provide a universal procedure. It proposes a way to read systems and intervene without separating architecture, observability and operational consequences.

Before acting, distinguish evidence from assumptions, identify genuinely shared dependencies and locate the failure domain or bottleneck that is shaping the service behaviour.

During an incident or a change, speed does not mean acting blindly. Limiting the blast radius, preserving a return path, observing the effects and verifying each step makes the intervention easier to govern and reduces the risk of amplifying the problem.

Resilience does not come from component duplication alone. It requires relevant signals, coherent health checks, explicit network flows, executable recovery procedures, verified backups and clear criteria for deciding when to continue, stop or roll back.

Availability, RPO, RTO, budget and complexity become meaningful only when they are connected to real dependencies and measurable actions. A broader solution is not necessarily a better one: it should satisfy the stated requirements without introducing costs and complexity that the context does not need.

Northbridge is fictional and the scenarios are deliberately bounded. They do not represent every variable found in a real environment, but they make decision criteria visible across changing technologies, tools and system sizes.