A hardened HTTP client for flaky providers
Every product I've worked on in fintech and identity eventually depends on someone else's API: a KYC vendor, a payment rail, a registry lookup. Those APIs fail in ways your own backend doesn't — rate limits, cold starts, partial outages, timeouts that sometimes mean "it worked".
On Kwikpik's verification platform I stopped letting each integration handle that on its own and wrote one shared client. These are the rules it enforces.
1. Auth is pluggable, not copy-pasted
Different providers want different things — a bearer token, an HMAC signature over the body, OAuth2 client credentials with a token that expires. The client takes an auth strategy, so a provider adapter declares which one it uses instead of reimplementing it.
type AuthStrategy =
| { kind: "bearer"; token: () => string }
| { kind: "hmac"; keyId: string; secret: string }
| { kind: "oauth2"; tokenUrl: string; clientId: string; clientSecret: string };The OAuth2 strategy caches its token and refreshes it a little before expiry, so a request never starts with a token that dies mid-flight.
2. Retry the right things, the right way
- 5xx and network errors: retry with exponential backoff plus jitter.
- 429: retry, but if the response says
Retry-After, wait exactly that long. The provider is telling you when it will accept you; guessing is rude and slower. - 4xx (except 429): never retry. A 422 is bad input or a bug. Retrying it three times just triples the noise and delays the error the user needs to see.
function shouldRetry(res: Response | null, attempt: number) {
if (attempt >= MAX_ATTEMPTS) return false;
if (!res) return true; // network error / timeout
if (res.status === 429) return true;
return res.status >= 500;
}3. Every retry carries the same idempotency key
A timeout is not a failure. It means "I don't know". If the first attempt actually succeeded and you retry without an idempotency key, you've created two verifications — or worse, two payouts.
So keys are generated per step, not per request, and reused across every retry of that step. The provider dedupes; we sleep at night.
4. Timeouts are mandatory
fetch will wait forever by default. Every call gets an AbortSignal.timeout(), and a timeout counts as a retryable failure and as a strike against the circuit breaker.
5. A circuit breaker for when it's really down
If a provider fails N times in a window, the breaker opens: calls fail immediately for a cooldown instead of each one waiting eight seconds to time out. After the cooldown it goes half-open and lets one probe through. Success closes it; failure opens it again.
The user-facing benefit is huge: during an outage, people get a fast "verification is temporarily unavailable — try again shortly" instead of a spinner that ends in the same message forty seconds later.
6. Errors are typed and redacted
Callers get a discriminated union — ProviderUnavailable, RateLimited, Rejected, Timeout — never a raw response. Logs record the provider, the step, the status and the attempt, and never the body of a request that contains a document number or a face.
Most of this is ten lines each. The value is that it's written once, tested once, and every new provider adapter gets all of it for free.