Our work
Systems in production for retail, airline, adtech and legaltech organisations.
Client case
Real-time event collection: −70% infrastructure cost, a pipeline still in production
A global retail leader · 9-month engagement · Tech lead of a team of 4
≈ 300 M
events per day, with peaks at 6,000 per second.
−70%
on the scope's infrastructure bill. Line-by-line FinOps forensics, then a Cloud Run → GKE migration.
The starting point.
Browsing-event collection relied on client-side tag management. Between ad-blockers, Apple's restrictions (ITP) and the weight of third-party JavaScript on page performance, a growing share of the signal was lost before it ever reached the analytics tools. Meanwhile, the cloud bill for the setup climbed with traffic.
The initial request was about signal loss. The diagnosis revealed two distinct problems: the signal, indeed, and an infrastructure cost whose cause was not the one everyone assumed.
What we built.
A first-party collection pipeline, in a real-time streaming architecture: a collection point under the client's domains, Kafka as the backbone, and a fan-out to consumers: analytics, personalisation, A/B testing, datalake.
Three structural choices, made as a team. A versioned event schema that acts as a contract between producers and consumers: an invalid event is rejected on the way in, not discovered three systems later. At-least-once delivery guarantees, with idempotent deduplication, exponential-backoff retries and replay queues (DLQ) so nothing is lost when a destination goes down. And consent treated as first-class data: the CMP state travels with every event and conditions its routing.
In production: on the order of 300 million events per day, with peaks at 6,000 events per second.
Where the −70% came from.
The bill did not come from a single line item. Analysing consumption line by line surfaced the classic accumulation of a system that grew fast: oversized logging, verbose and uncompressed payloads, caching missing where it would have mattered, memory leaks in some services. We fixed it item by item, before touching the platform.
That left the execution model. Cloud Run is simple and fitting at low volume, but its per-request billing becomes structurally unfavourable against a sustained load of several thousand events per second. The migration to GKE (bin-packing, autoscaling, resource right-sizing) aligned cost with real load.
The two effects combined: −70% on the scope's infrastructure bill.
What the system became.
The pipeline is still in production. The sign that counts: adding a new event type became a configuration operation, covered by automated regression tests, rather than a development project. That is how you recognise a durable foundation.
Proof by product
We publish digitalyser.io
digitalyser.io is our SaaS platform for digital visibility: SEO audits, analytics, e-reputation, AI-driven content. We design, operate and evolve it with the same engineering practices we apply for our clients. A living reference: a product in production, with real users.
Visit digitalyser.io ↗Sectors
They trusted us
- Decathlon
- Air Austral
- Mediarithmics
- NeoNotario
A data, AI or infrastructure project?
Let's talk. Thirty minutes, no strings attached. At the very least, you leave with an engineer's take on your problem.
Book a call