Skip to content

Why Open Banking connections break — and how to monitor them

Updated · MVP Payments

A bank connection is not something a PISP builds once. It is something that works until the bank changes — and banks change constantly. Ask any team that has run Open Banking payments in production and you will hear the same story: a connection that processed payments yesterday fails today, and nothing on the PISP’s side has moved.

This guide sets out why that happens and what to monitor so that you find out from your dashboards rather than from your merchants.

What the rules promise

The PSD2 regulatory technical standards give third parties real protections. A bank’s dedicated interface must offer the same availability and performance as its own customer channels, the bank must publish quarterly statistics to prove it, and changes to the interface’s technical specification should be communicated at least three months ahead, except in an emergency.

In practice, notice periods cover formal specification changes. They do not cover the much larger category of behaviour a specification never described.

How connections actually fail

Certificates

The most common cause, and the most avoidable. A PISP’s own QWAC or QSealC expires, and every payment at every bank stops at once. Less obviously, banks rotate their certificates and certificate chains too, and a third party that has pinned the old chain starts failing its TLS handshakes.

Changes beneath the specification

A field the specification marks optional becomes mandatory. A header that was ignored is now validated. An error that came back as one status code now comes back as another. None of this is a specification change, so none of it is announced — and any of it can take a connection down.

Authentication journey redesigns

Banks rework their strong customer authentication flows regularly: a new app, a new redirect domain, a different hand-off on mobile. The API contract is untouched, but payers begin dropping out at a step that used to work. This failure is especially hard to see, because every API call still succeeds.

Shared IT providers

In markets such as Germany and Austria, hundreds of banks reach third parties through a few shared banking IT providers. One release by one provider changes the behaviour of every bank behind it on the same day. It is the largest single cause of wide, sudden incidents.

Version upgrades and retirements

Standards move on and banks follow at their own pace, retiring old API versions on their own schedules. A PISP with many connections is always mid-migration somewhere.

Sandbox drift

Bank sandboxes are separate systems that often lag or differ from production. A change can pass every sandbox test and fail live — or fail in the sandbox while production is fine, sending a team after a problem that does not exist.

Maintenance windows and limits

Planned maintenance, rate limits and overnight batch windows all produce failures that are expected. If monitoring cannot tell them apart from real incidents, it will be ignored by the time a real one arrives.

What to monitor

Success rate per bank, not overall. An aggregate success rate hides a failing bank behind healthy ones. Measure every connection separately and judge each against its own baseline.

Stage of failure. A payment can fail at creation, at authentication hand-off, at the bank’s execution or at status retrieval. The stage points straight at the cause.

Bank-side versus payer-side outcomes. A payer who cancels is not an incident. A monitoring model that cannot separate cancellations from technical failures will either cry wolf or go quiet.

Latency. Rising response times often precede failures, particularly ahead of a bank-side release.

Status progression. Payments that stall — accepted but never confirmed — are a failure mode of their own. It helps to know each bank’s normal ceiling: some never report beyond “accepted”, and that is not a fault.

Certificate expiry. Yours and, where visible, the bank’s. Thirty days’ warning is comfortable. Seven is a deadline.

State changes, not single errors. Alert when a connection moves from healthy to degraded to down on sustained evidence. One failed call is noise.

The part monitoring cannot do

Detection is half the work. The other half is diagnosis, and diagnosis depends on knowing the bank: what changed the last time this happened, which field this bank is particular about, whether its sandbox can be trusted. In most teams that knowledge sits with one or two engineers — and leaves with them.

This is where MVP Payments uses AI. Every connection’s health is re-evaluated every five minutes, and every connected bank has a dedicated AI agent that holds a standing record of that bank and a dated journal of every past investigation. When a connection degrades, the agent starts from that history and prepares the diagnosis and the change; an engineer reviews and releases it. Read how our connectivity monitoring works.

Why it is easier with direct connections

A PISP can only monitor what it can see. Behind an aggregator, a failing bank and a failing aggregator look the same, and the raw bank responses that would settle the question belong to someone else. A direct connection — to the bank itself, or to the national hub the bank uses — returns the bank’s actual behaviour, which is the only reliable basis for saying what changed.

If maintaining that estate is not where you want your engineers, compare the alternatives in build versus buy for Open Banking connectivity.