Risk Management in Dubai: Rehearsed Plans | Ilia Arestov

Risk Management and Business Continuity

Risk management is easy to confuse with the spreadsheet somebody fills in once a year, shortly before the auditor arrives. I am called in when a working version is needed: a short list of the threats that could halt the business, each with an owner and a rehearsed response plan. Rehearsed means it has been run at least once, not merely written.

What people arrive with

Risk register and business continuity plan: critical points with named owners

The first conversation usually runs off somebody else’s paperwork. A bank or a large client has sent a fifty-question continuity questionnaire and given ten days. Then it asks for the tolerable downtime of billing, who may declare an incident, and when a backup was last restored, and the filling in stops. The problem is rarely disorder: backups run every night, and a risk register does exist, forty rows marked “medium” with not one name against them. It is simply that none of those numbers has ever been said out loud.

Risk management starts with the register

The register is built from the money, not from a catalogue of threats: what stops earning revenue tomorrow if a system, a person or a supplier disappears today. In a company of a hundred people there are rarely more than ten critical points, and half of them are not technical at all: they are the one employee who keeps the calculations in their head.

ISO 31000 provides the skeleton, but the standard is a way of writing things down, not the goal. The test for a row is simple: if nobody can say what happens to it next Monday and whose signature ends it, the row was written for nothing.

  • Dependencies mapped: which functions are critical and where the point of failure is single.
  • Impact estimated in money and hours of downtime, not in scores.
  • The owner of a risk is a person, not a department.
  • A decision on every item: close it, insure it, or accept it deliberately.

Under DIFC, ADGM or the UAE Central Bank requirements the register and the plans are written to pass inspection. Configuring the perimeter itself stays outside this engagement: that is information security, with its own scope and its own timeline.

Continuity and recovery

A continuity plan starts with two numbers for each critical function: how long it can be down and how much data may be lost. The business states them, not IT. A target of “back in four hours” more often means that is what the current architecture allows, not that four hours of downtime is acceptable.

Those numbers then turn into a design, and the conversion is arithmetic: four hours means the image has to come up in two, because detection, the decision and the check eat the other half. Full compromise of the environment is costed separately: a copy reachable with the same key as production restores nothing. ISO 22301 is the reference; a copy that has never been restored does not count as a backup, so testing it is part of the engagement. Stated availability and tolerable downtime are two different promises: 99.9 % on Monolith Plus is roughly eight hours a year, and that still has to be split between maintenance windows and real failures.

Incident response

During an incident there is no time to work out who is allowed to do what. So the plan answers the boring questions in advance: who declares an incident, who may stop a service, who speaks to the client and to the regulator, and which channel people use when the mail system is down because of the incident itself.

The structure comes from NIST SP 800-61: detection, triage, containment, eradication, recovery, review. Each scenario gets a one-page or two-page instruction: ransomware, a failed supplier, a data leak. The signal itself comes from monitoring, which is covered in A SIEM on an open core.

Rehearsals: a plan that has been run

One thing separates a working programme from a folder of documents: it gets run, and you watch what breaks.

  • Tabletop walk-through: everyone says what they do in the first fifteen minutes.
  • A real restore from backup on a separate environment, timed with a stopwatch.
  • A supplier fails: the payment provider closes the account for a day and money still has to be taken tomorrow.
  • A key person is missing: whoever closes the month is unreachable for three days and the filing is due Friday.

Stack

The table holds only what serves one of the two numbers: the recovery deadline or the depth of data loss.

What it provesTools
A copy that gets restoredProxmox Backup Server, OpenZFS
The environment a rehearsal runs onProxmox VE
Downtime measured and alerted onUptime Kuma, Grafana
Switchover written down as a playbookAnsible

What I do not do

  • No ISO 22301 certificate comes out of this engagement. A rehearsal report does: dates, measured recovery times, and a list of what broke during the run.
  • I do not write a register from a description given in a meeting: without access to the processes you get a document, not protection.
  • I do not replace an insurance broker: I take the decision to insure as far as an expected-loss figure, and a broker places the policy.
  • No backup vendor or insurer pays me for a mention; licences you buy directly and in your own name.

How the engagement runs

The first thing the engagement produces is two numbers for one function. At the first meeting the most revenue-heavy function is taken and its tolerable downtime and tolerable data loss are said out loud; that is usually where two managers turn out to have different figures in mind. Two weeks then go on the processes and the dependencies, and the result is a map of the critical functions carrying those same numbers and whatever stops them being met.

The main part runs eight to fourteen weeks for a company of 30 to 150 people: the register, the continuity plan, the incident instructions, monitoring and backup configured, and handover to your team. It closes with rehearsals and the revisions they produce.

The project is priced as a fixed sum, with $30,000 (AED 110,100) as the reference; the exact figure is fixed after the survey.

Technical defence of the perimeter is information security. Risks inside a three-year plan belong to strategic planning. Liquidity and currency exposure belong to financial management. For an independent view of the whole stack, see IT consulting.

Frequently Asked Questions

When can the first rehearsals happen?

The tabletop walk-through lands in week five, a full restore on a separate environment between week ten and week fourteen. Getting the participants into one room for two hours is usually harder than configuring the backups.

How is this different from information security?

Security covers the technical perimeter: access, devices, encryption, monitoring. Risk is wider: a key employee leaving, a supplier failing, a provider account lost. Cyber risk is one row in the register.

Who has to be interviewed for the register to be more than paper?

The people who run the process by hand: the accountant names the most expensive day of the month, the administrator gives the time it took to bring the last failed server back.

Who keeps the register alive once the project ends?

One person from your own management, named before the start. The register then runs on a calendar: a review every quarter and one after every incident. What stays is the tested RTO and RPO, the instructions, monitoring and backup configured, and the rehearsal report.


Ready to Get Started?

Name the function whose outage costs you most: one conversation shows whether you need the full programme or just two or three fixes. Talk through that first function: the consultation is free.