
Technology
Tenant-Aware Scaling for a Growing SaaS Product
Tenant-aware queues separate imports and background processing from everyday product use, with workload limits, recoverable jobs and controlled releases.
Explore the solutionSolution design
Large jobs stop setting everyone's pace
The application separates lengthy imports and exports from interactive product use. Tenant-aware limits protect shared capacity, while queued jobs retain status, retry history and recovery information. Releases account for records already in flight. The operating view connects customer-facing delays to the workload responsible, giving engineering and support a specific starting point for investigation.
Customers can keep using the product while imports and exports run in the background. Each operation has a status, a retry history and a way to recover. Engineering can see which workload needs capacity before increasing resources across the entire application.
- Interactive requests have protected capacity
- Interrupted imports resume without duplicate records
- Support can explain the state of a job
- Business context
- A B2B software provider serving accounts with different data volumes, feature entitlements and integration needs.
- Core capability
- Software Development
The busiest account is not always the biggest
A customer importing a large dataset can put more pressure on the product than many people browsing it. The import uses database connections, processor time and storage bandwidth that ordinary requests also need. User counts alone miss this pattern. The relevant question is what each account asks the system to do, and when.
Product requests add another kind of pressure. A custom report or integration can introduce expensive work into code shared by every account. Before accepting that work, the team decides whether it belongs in the main product, a configuration option or a separate integration. That decision affects both performance and the cost of maintaining the feature.
Trace a slow operation before choosing a fix
Start with an import and the everyday screens that slow down around it. Trace the request through database queries, external calls and background workers. Include work that continues after the browser receives a response. A quick acknowledgement can hide a queue that takes hours to clear or repeatedly starts the same failing batch.
Record the account, operation and attempt identifiers alongside timings and resource use. Keep customer file contents out of routine logs. These traces help distinguish a missing index from a worker shortage, or an expensive calculation from a slow supplier API. Each requires a different fix; increasing the application fleet can leave the actual constraint untouched.
Solution scope
- Request, query and background-job tracing
- Tenant permissions and workload limits
- Modular application boundaries
- Queued imports with row-level error reports
- Database migration and rollback procedures
Separate responsibilities before splitting services
Identity, subscription rules, product records and reporting need clear boundaries in the code. Each module owns its rules and exposes a small interface. They can still run as one application. A separate service is useful when a workload needs independent capacity or availability, and the team is ready to operate another deployment and failure boundary.
Lengthy imports and exports run through a queue. Interactive requests validate the operation, record it and return a reference. Workers take bounded batches, leaving capacity for normal use. Database migrations add compatible fields before retiring old ones, so running jobs and successive application versions can coexist while a release moves through production.
Show what happened to the file
The upload first passes permission, file-size and format checks. The application creates an operation record before queuing any processing. A worker then validates rows and separates rejected entries from accepted changes. Progress shows processed records and outstanding work. The error report names the affected rows and what the customer needs to correct.
A stable operation reference follows every processing attempt. Accepted changes are repeat-safe, so recovering a failed batch does not insert the same records again. Partial completion remains visible. Support can inspect the stage, last attempt and failure category without querying production tables or asking an engineer to search unrelated logs on the customer's behalf.
The operational flow
Check the request
Resolve account permissions and validate the file before accepting it.
Create the operation
Store a reference that the customer, worker and support team can use.
Queue the batches
Assign heavy processing to workers within the account's capacity limits.
Apply and record
Save accepted changes and row errors with a repeat-safe batch history.
Finish or recover
Show the result, or route the failed stage to someone who can resolve it.
Move one troublesome operation first
Begin with tracing and a mixed-workload test that reproduces the contention. Move one operation behind the queue while keeping the customer journey recognizable. Release it to a limited account group, compare its records with the existing process and check how support handles a failure. This exposes missing status information before the pattern spreads.
Rehearse migration and recovery using a realistic dataset. Rollback instructions must account for records already changed, not only the application version. Check that old and new workers can finish their jobs during deployment. Move the next workload when its measurements justify the change, rather than treating the queue as a reason to redesign every feature.
Give support an answer before the next ticket arrives
Support needs to know whether a job is waiting, making progress or awaiting intervention. The operating guide explains those states, the checks staff can make and when engineering takes over. It also distinguishes a safe retry from a correction that needs a developer. A visible retry button should not grant permission to repeat any operation.
Engineering maintains processing rules and alerts. Product owns entitlement changes and the customer experience during interruptions. Infrastructure owners review capacity and cost. Alert on the oldest waiting work as well as queue size: a small queue can still contain a stranded customer. Recurring failures go into product planning with their cause and recovery effort attached.
Test the quiet screens while the heavy work runs
Run ordinary transactions alongside large imports, using a realistic mix of accounts. Compare response times, failures and queue age across those workloads. Check whether an account near its limit gets an understandable explanation and whether other accounts remain responsive. Include recovery runs, because a design that behaves well only on the first attempt is incomplete.
Relate infrastructure spending to completed operations and their sizes. A large export naturally costs more than a simple page view. Keep that distinction in capacity reviews, alongside release incidents and support effort. The useful next investment may be a better query, a different batch size or clearer recovery controls rather than another server.
The choices behind the solution
Keep one deployment where it works
Use modules before adding independently deployed services.
A queue can isolate heavy processing without multiplying service-to-service dependencies.
Make job state part of the product
Expose progress, partial completion and actionable row errors.
Customers otherwise retry blindly, adding load and creating more support work.
Limit competing workloads
Apply account and operation limits separately from subscription entitlements.
Permission to use a feature does not mean unlimited access to shared resources.
How the solution is evaluated
These measures define the evaluation criteria for the workflow, its controls and the quality of completed tasks.
Response time during bulk processing
Measure: Compare the same interactive journeys while different account workloads run.
Success criteria: Ordinary product use stays within the agreed service limits.
Age of unfinished jobs
Measure: Track the oldest waiting and failed operations, including repeat attempts.
Success criteria: Support can identify stranded work and follow a documented recovery route.
Resources per completed operation
Measure: Group compute and database consumption by operation type and workload size.
Success criteria: Capacity changes address the work consuming resources rather than an unexplained total.
The architecture has to serve two different rhythms: people waiting for a screen and jobs working through large amounts of data. Giving each its own limits, status and recovery path makes that distinction manageable. It also gives the team a practical way to extend the product without turning each new workload into a system-wide performance problem.
