Lead-generation SaaS
Removing a SaaS platform's scalability ceiling
Ten concurrent users in the editor degraded the entire product to multi-minute waits. We isolated that workload first, then decomposed the rest — and the customer base grew sixfold without the architecture becoming the limit.
- Period
- 2018 – 2019 · 13 months
- What we delivered
- Architecture audit, monolith decomposition, AWS migration, CI/CD and observability
- Setting
- Backend lead within a 7-person product team
Client and project names are withheld under confidentiality. Sector, scale and outcomes are reported as delivered.
300 → ~2,000
Customer accounts during the engagement
Minutes → none
Perceptible slowdown under concurrent editing
4 services
Domain-driven services from one monolith
VPS → AWS
Infrastructure migration, independently scalable
Context
A self-service SaaS product letting website owners design, target and deploy on-site popups without writing code. Customers configure design, content and display rules in a web dashboard, test them, then embed a single JavaScript file on their site; the platform evaluates the configured conditions against each visitor and renders accordingly, with built-in analytics.
The targeting engine was the product’s core differentiator, supporting close to 50 combinable display conditions — visitor country, referring site, prior visits to a given page, dwell time on a specific page, even whether a particular image on a particular page had entered the viewport.
Problem
The product had grown quickly from MVP into a monolithic application on a single VPS, with no automated tests, no linting and no CI/CD. Weekly releases were manual.
One bottleneck was existential. Because the editor offered live preview, every change a user made regenerated the popup’s JavaScript bundle — and users iterated heavily. With only 10–15 concurrent users in the editor, the entire platform degraded, with wait times reaching several minutes, despite modest overall traffic.
That is not a performance problem. That is a ceiling on how many customers the business can have.
Approach
Audit before architecture
We audited the existing system and produced a formal written report on structural weaknesses and technical debt, then translated it into a prioritised, sequenced implementation plan that could be absorbed alongside the product roadmap rather than blocking it. A decomposition that requires a feature freeze does not get approved, and should not.
Isolate the dominant load first
We led the decomposition into domain-driven microservices — users, popups, analytics, reporting — defining each service’s bounded context, explicitly scoping what it owned and, equally importantly, what it did not.
The sequencing mattered more than the target shape. The JavaScript-generation workload moved into a dedicated popup service first, containing the platform’s dominant load in one place instead of letting it degrade every other capability. We then optimised that service specifically — Redis caching, queue-based processing, right-sized infrastructure — and made it independently scalable both horizontally and vertically as demand grew.
The analytics service was built on Elasticsearch, giving the reporting domain a query engine suited to its access patterns rather than forcing it onto the transactional database.
Infrastructure and delivery
- Drove the migration from a single VPS to AWS, selecting the services and defining how each would be used: EC2, SQS, RDS, ElastiCache, RabbitMQ, S3 and load balancing.
- Introduced automated testing and linting to a codebase that had neither, built into the pipeline as quality gates.
- Designed the Docker containerisation and CI/CD pipeline from scratch, replacing manual releases with an automated build-and-deploy verified through health checks and log inspection, with defined rollback on failure.
- Established observability — service health checks, centralised error-log aggregation across every service, and alerting on critical failures — giving the team its first proactive view of production problems.
Development and team
- Built reference implementations demonstrating the target architecture in production code, so the standard existed as something to follow rather than an abstract guideline.
- Refactored payments into a provider-agnostic abstraction. The platform supported one provider; we integrated a second and, in doing so, restructured the payment layer behind a common interface so further providers could be added without touching business logic.
- Owned code review across the backend team and mentored the junior developer through pair programming and review feedback.
- Contributed to the React front end where it met the backend — dispatching editor changes over the message queue and applying the resulting updates to the live preview as queue responses arrived.
Outcome
- The platform’s central scalability ceiling was removed. Concurrent editing that previously pushed the whole system to multi-minute waits ran with no perceptible slowdown, and the popup service could be scaled independently as load increased.
- Supported growth from roughly 300 to about 2,000 customer accounts during the engagement, without the architecture becoming the limiting factor.
- Established the platform’s entire quality and delivery foundation — tests, linting, containerisation, CI/CD, health-checked deployments and rollback — where none had existed.
- Made payment integration a configuration concern rather than a development project, unlocking new providers and markets cheaply.
- Left the team with clear service boundaries, documented architectural direction and working reference implementations.
Related work
Enterprise retail e-commerce
Taking an enterprise commerce platform from 1,000 daily errors to double digits
A composable-commerce platform for a large optical retail chain was logging around a thousand errors a day, timing out on discounts and charging the wrong amounts. We worked the defect classes down in order.
~1,000 → double digitsError log entries per day
Read the case study →
Regulated consumer products · D2C
A regulated direct-to-consumer storefront, live and transacting in 45 days
We won the engagement on a commitment competitors would not make — a custom storefront taking live card payments within 45 days, under a 100% uptime guarantee. Both were met.
45 daysFrom zero to live card payments
Read the case study →
Building automation · IoT
A building automation platform where a device contract is defined once
A KNX hardware manufacturer needed a software platform to control its own devices and any third-party KNX device. We defined the architecture and built the logic engine, the visual automation editor and the tablet client.
One contractNode schema shared by backend, admin panel and mobile
Read the case study →