Sanitized accounts of real engagements. Company, customer and internal project details are removed. The focus is the diagnostic process and the architectural reasoning, since that is the part that transfers.

No performance metrics are quoted where the exact figures are not known. A plausible-sounding number is worse than no number when you are trying to learn from it.

A commit pipeline that was too slow

Stack: Scala, Akka/Pekko, Kafka, Cassandra, Docker

Problem

Commit rate on a production event-driven platform was too slow to keep up with the workload arriving on the topic.

Investigation

The investigation did not point at a single component. Four contributing factors stacked up on the same path:

  • synchronous thread blocking in the application
  • inefficient Kafka topic configuration
  • a single Kafka partition, which capped parallelism
  • a Cassandra cache pattern of quick reads, writes and deletes
Application
    ↓
Synchronous thread blocking
    ↓
Kafka
    ↓
Single partition
    ↓
Cassandra
    ↓
Cache read/write/delete

The useful conclusion was not any individual fix. It was that the bottleneck crossed several system layers, so no single component's metrics would have revealed it. The single partition in particular looked like a configuration detail until it was placed next to the blocking calls upstream of it.


Kafka lag traced back to Cassandra tombstones

Stack: Scala, Akka/Pekko, Kafka, Cassandra, Docker

Problem

Kafka consumer lag kept increasing. The natural assumption was that the consumers were too slow, or that the producers were too fast.

Investigation

Consumer throughput was not the origin of the problem. Downstream Cassandra read performance had degraded because expired rows were being deleted individually, producing a large number of tombstones. Slower reads reduced processing throughput, which is what consumer lag was actually measuring.

Kafka lag
    ↓
Processing throughput
    ↓
Cassandra read performance
    ↓
Tombstones
    ↓
Data lifecycle/deletion strategy
    ↓
Data model redesign

Resolution

The data model was changed to a week-based partitioning and data lifecycle strategy. Instead of deleting expired rows one at a time, the previous week's data could be removed with the primary-key operation that matched the partitioning.

The general lesson: when a lagging queue is the visible symptom, look at what the consumers are actually waiting on before tuning the consumers.


A reporting system that failed at scale

Stack: Java, MySQL, BIRT, Servlets

Problem

An MIS and reporting system worked fine with small datasets and became effectively unusable in production with large customer datasets.

Resolution

  • restricted reporting data to a one-month window instead of querying indefinitely
  • made report generation asynchronous, so a user could submit a request and return later to collect the generated file
  • added MySQL indexing
  • divided database and data handling for large customers

Architectural lesson

Scaling this system required changing the data-access model and the user workflow, not merely optimising queries. Making report generation asynchronous changed the shape of the problem rather than shaving time off it, and the one-month window was a data-model boundary rather than a filter.