Part 3 · 10 chapters · ~55 min

Security: threat models, not checklists

A checklist tells you what other people’s systems needed. A threat model tells you what yours needs, and it takes forty minutes with a diagram and six prompts. This part walks the trust boundaries of the bank built in the CBA module, and spends its longest chapter on the insider threat, because in financial services that is consistently where the loss actually comes from.

28

Threat modelling in forty minutes

the question

“Is this design secure?” is unanswerable. “What can go wrong, and what stops it?” is a forty-minute meeting.

the four questions, which is the whole method
  1. What are we building? A diagram with trust boundaries drawn on it. Not the architecture diagram: one that shows where data crosses from less trusted to more trusted.
  2. What can go wrong? Walk each boundary with a prompt list. STRIDE is the common one: spoofing, tampering, repudiation, information disclosure, denial of service, elevation of privilege.
  3. What are we going to do about it? Per threat: mitigate, transfer, accept, or eliminate. Accept is a legitimate answer when recorded.
  4. Did we do a good job? Revisit when the design changes, and after any incident.
worked numbers
walking the boundary between the app and the posting API:

  Spoofing       can a caller pretend to be loans?  → mTLS, Part 14
  Tampering     can entries be altered in flight? → TLS, and Σ=0 server-side
  Repudiation  can an action be denied?      → append-only audit log
  Info disclosure can it read other accounts?   → scope by account kind
  DoS             can one caller starve others?   → per-identity rate limits
  Elevation     can it grant itself overdraft?   → Part 18: floor verified

six prompts, one boundary, and most of the CBA design justified.
why this beats a checklist
A checklist tells you what other people's systems needed. A threat model tells you what yours needs. The CBA design already contains the answers, and walking the boundaries is how you discover whether they are actually there. Forty minutes with a diagram and six prompts finds more than a day with a compliance questionnaire.
29

The trust boundaries in our bank

BoundaryCrossingPrimary control
Internet → edgeUntrusted, anyoneWAF, TLS, rate limiting, authentication
Customer → their own dataAuthenticated, still untrustedAuthorisation on every object, not just the endpoint
Service → serviceTrusted-ishmTLS identity plus scope by account kind
Service → ledgerThe critical oneServer-side invariant validation. A caller bug cannot create money
Staff → customer dataThe insider boundaryFour-eyes, rate limits per agent, full audit
Provider → us (webhooks)Untrusted, looks trustedSignature over raw bytes, replay bound, amount verified
Us → providerOutboundOur reference, idempotency, allowlisted egress
Region → regionLegal boundaryNo route, no schema field, org policy
the boundary people forget
Staff to customer data. Every other boundary has a natural adversary and gets attention. This one is crossed thousands of times a day by people who are supposed to, which is why Part 18 made restricting easy and unrestricting hard, and why Part 9 monitors the audit log itself. The webhook boundary is second: it looks internal because it arrives over an established integration, and it is a public endpoint that moves money.
30

Authentication, authorisation, and the gap between

worked numbers
authentication: who are you?
authorisation: what may you do?

the gap where bugs live:
  authenticated, therefore authorised

the classic form, which is still the most common serious bug:
  GET /accounts/{id}/statement
  auth check: is this a valid session?         yes
  authz check: is this account THEIRS?       missing

IDOR: insecure direct object reference. it is boring, it is
everywhere, and it is the top cause of real data breaches.
the defences, in order of durability
  1. Authorise the object, not the endpoint. Every handler that takes an id must check ownership of that id, and it must be structurally impossible to forget.
  2. Scope queries at the data layer. If the repository takes the caller's identity and filters by it, forgetting the check does not leak data. Make the safe path the default path.
  3. Unguessable identifiers are defence in depth, not a control. A UUID does not authorise anything; it just makes enumeration harder.
  4. Test for it. An automated test that tries to read another customer's data and expects a 403, on every endpoint that takes an id.
31

Secrets: the lifecycle, not the store

the five stages, and where each goes wrong
  1. Creation. Generated with a CSPRNG, never chosen, never in a ticket or a chat message.
  2. Distribution. Fetched at runtime by workload identity. Never in an image, a repo, an env file in git, or a CI log.
  3. Rotation. Scheduled and tested. The application must handle a credential changing underneath it, which is an application property rather than a vault feature.
  4. Revocation. Possible within minutes, which requires knowing every consumer of every secret. Most organisations cannot answer this.
  5. Audit. Every access logged, because anomalous secret access is an early compromise signal.
the uncomfortable question
"If a laptop were stolen today, what would we have to rotate, and how long would it take?" If the answer is unknown or measured in weeks, the secret management is inadequate regardless of which vault is in use. Rotation capability is the control; storage is the prerequisite. The cloud module made the same point and it is worth repeating because most teams store well and rotate never.
32

Cryptography choices a non-specialist must get right

NeedUseNever
Password storageArgon2id, or bcryptMD5, SHA-256 alone, any fast hash
Data at restAES-256-GCM via a KMSYour own key derivation
Data in transitTLS 1.2 or 1.3Anything you implement
Message authenticationHMAC-SHA-256Comparing with ==. Constant time
Random valuescrypto.randomBytes, CSPRNGMath.random(), ever
SignaturesEd25519 or ECDSAHomegrown schemes
TokensJWT with care, or opaquealg: none, unverified claims
the four rules for people who are not cryptographers
  1. Never implement a primitive. Use a vetted library, and use the high-level interface it offers rather than the building blocks.
  2. Constant-time comparison for anything secret. Part 6 of the CBA module made this point for webhook signatures: a fast-failing comparison leaks the correct prefix byte by byte.
  3. Encryption is not authentication. Use an AEAD mode such as GCM, so tampering is detected rather than silently decrypted into garbage.
  4. Key management is the hard part. The algorithm is almost never the weakness; where the key lives, who can read it and how it rotates is.
33

The insider threat, which is the real one

worked numbers
external attacker: must find a vulnerability, exploit it,
  escalate, move laterally, and exfiltrate, all undetected

insider: already authenticated, already authorised,
  already inside, and their actions look legitimate

in financial services, insider-enabled loss consistently
outweighs external intrusion. the controls are different.
the controls that actually work
  1. Least privilege, enforced. Not "we trust our people" but "nobody needs production data access by default, and access is time-boxed and approved".
  2. Four-eyes on the dangerous direction. Part 18: one agent may restrict, two are required to unrestrict. Asymmetric by the direction of risk.
  3. Rate limits on staff actions. An agent viewing 400 customer records in an hour is an anomaly regardless of whether each view was individually permitted.
  4. Audit the audit log. Part 9 monitors privileged actions for anomalies, and the log lives in an account production cannot write to.
  5. Separation of duties. Whoever can deploy should not also be able to approve a manual adjustment.
  6. Make the legitimate path easy. If getting proper access takes three days, people share credentials, and you have lost attribution entirely.
the cultural point
Insider controls are not an accusation. Framed as "we protect you from being suspected", four-eyes and audit become something engineers want: if you cannot act alone, you cannot be blamed alone. Teams that frame it as distrust get resistance; teams that frame it as protection get compliance.
34

Supply chain: dependencies and build integrity

the attack surface, in order of likelihood
  1. A vulnerable dependency. The common case, and largely a patching-speed problem rather than a detection one.
  2. A malicious dependency. Typosquatting, or a maintained package that changes hands. Growing steadily.
  3. A compromised build. The pipeline has credentials to everything and is often the least protected system in the estate.
  4. A compromised artefact registry. Signed images and verified provenance exist for this reason.
ControlWhat it gives you
LockfilesThe same versions every build. Table stakes
SBOMAn inventory, so "are we affected?" is a query rather than a week
Automated scanningContinuous, with a patch SLA. The SLA is the control, not the scanner
Signed artefactsProvenance from source to running image
Pinned base imagesBy digest, not by tag. A tag moves
Restricted pipeline credentialsScoped and short-lived, not a long-lived admin token
the question that reveals maturity
"A critical CVE drops in a common library at 4pm on a Friday. What happens?" A mature answer is: the SBOM says within minutes whether we use it and where, the patch SLA says how fast, and the deployment pipeline can ship it safely today. An immature answer involves grepping repositories.
35

Security in the pipeline, not after it

what belongs in CI, ordered by value per minute of build time
  1. Secret scanning. Blocks a committed credential before it ever reaches a remote. Highest value, lowest cost.
  2. Dependency scanning with a policy: fail on critical, warn on high.
  3. Static analysis tuned to near-zero false positives. A noisy scanner gets ignored and then disabled, which is worse than not having one.
  4. IaC scanning. Catches a public bucket or an open security group before it exists.
  5. Authorisation tests, per chapter 30: every id-taking endpoint tested for cross-customer access.
  6. Not: a full DAST scan on every commit. Too slow. Run it nightly against staging.
the principle
Shift left, but only what is fast and precise. A security gate that adds fifteen minutes to every build and produces false positives will be bypassed within a month, and the bypass becomes permanent. Two fast precise checks that run every time beat six thorough ones that get disabled.
36

Incident response when it is an attack

what differs from an ordinary incident
  1. The adversary is watching. They may see your Slack, your deploys, and your investigation. Move to an out-of-band channel immediately.
  2. Evidence matters. Snapshot before you remediate. Terminating the compromised instance destroys the forensics.
  3. Containment before eradication. Isolate, revoke credentials, and cut access before hunting for root cause.
  4. Assume broader compromise. One compromised credential implies others. Rotate the blast radius, not just the known one.
  5. Legal and regulatory obligations start immediately. Breach notification clocks in many jurisdictions begin at discovery, and they are measured in hours.
  6. Do not tip off. If it is an insider, a visible investigation prompts destruction of evidence.
the preparation that pays
An out-of-band communication channel, established before you need it, with the people who need to be in it already there. Setting one up during an incident, while assuming the primary channel is compromised, is both slow and conspicuous. Ten minutes of preparation, and it is one of the highest-return items in a security programme.
37

The security review, as a checklist

for a service on the money path
  1. A threat model exists, with trust boundaries drawn and STRIDE walked per boundary.
  2. Every object access is authorised, not just every endpoint, and there is a test that proves it.
  3. Service identity is cryptographic and scoped by resource, not only by method.
  4. Secrets are fetched at runtime, rotate on a schedule, and rotation has been executed at least once.
  5. No custom cryptography. Constant-time comparison on every secret comparison.
  6. Inbound webhooks verify a signature over raw bytes, bound replay, deduplicate, and re-verify the amount.
  7. Staff access is least-privilege, time-boxed, four-eyed on the dangerous direction, and rate-limited.
  8. Audit logs are immutable, in a separate account, and include data access.
  9. Dependencies are scanned with a patch SLA, and an SBOM exists.
  10. An out-of-band incident channel exists and the right people are already in it.
the honest framing
None of this makes a system secure. It makes a system one where the common failures have been considered and the expensive ones are structurally harder. Security is a property you maintain, not a state you reach, which is the same thing Part 0 said about every other property in this module.