Principal Engineer & Independent Technical Consultant
Ayush Vatsyayan
Distributed Systems · Event-Driven Architecture · Data Platforms · Performance & Scalability
I help engineering teams design, troubleshoot and scale distributed systems, event-driven platforms and data-intensive applications.
Problems I help solve
Most of the engagements I take on start as a symptom somewhere in the stack and a bottleneck somewhere else. The work is tracing the path between them.
Independent architecture review
A second opinion on an existing or proposed architecture, before or after a significant investment.
Modernize Legacy Technology
Modernize legacy technology and identify practical analytics opportunities acros engineering and operations.
Difficult production problems
A live incident or recurring failure where the real bottleneck has not been identified yet.
Slow or unreliable distributed systems
Throughput is not keeping up, scaling is unpredictable, or behaviour under load is hard to reason about.
Kafka consumer lag and event pipelines
Lag keeps growing, processing throughput has plateaued, or partitioning and topic configuration need a second look.
Cassandra performance and data modelling
Queries degrade at scale and read performance are hurting, or a data model needs reviewing.
Selected case studies
Sanitized accounts of real engagements, focused on the diagnostic process and the architectural reasoning rather than company specifics or invented metrics.
-
A commit pipeline that was too slow
A single-partitioned Kafka topic, synchronous blocking in the application thread and a cache read/write/delete pattern in Cassandra all contributed to the same problem.
-
Kafka lag traced back to Cassandra tombstones
Rising consumer lag turned out to be a data lifecycle problem. Expiring rows individually produced enough tombstones to degrade reads feeding the pipeline.
-
A reporting system that failed at scale
Fixing it required changing the data-access model and the user workflow, not optimising the existing queries.
Engagement model
I am currently running a focused review session as a first, low-commitment way to work together. Have a distributed-system performance or scalability problem you cannot pin down?
Distributed Systems Architecture & Technical Review
- 60–90 minute technical session
- Review of the architecture, the problem and available evidence
- Discussion of relevant logs, code and design
- Identification of the likely bottleneck and its contributing factors
- Written findings
- Prioritized recommendations and follow-up Q&A
Initial rate: ₹5,000
This is only a diagnosis and review engagement.
Background
18 years of production systems
Senior and principal engineering across telecom, analytics, machine learning, big data and enterprise platforms, mostly as an individual contributor.
Event-driven platforms at scale
Designed and led an event-driven ETL platform with high-throughput aggregation pipelines, and distributed event-streaming engines modelling complex Layer-3 network topologies. Scala, Akka/Pekko, Kafka, Cassandra, REST APIs, Docker, Kubernetes.
Data platforms and ML
Moved a text analytics platform from Excel and small local datasets onto Spark, PySpark, Hadoop and Scala. Earlier work across Hadoop, Hive, Spark, Cassandra and telecom analytics at Nokia and Alcatel-Lucent.
Working with engineering teams
Architecture and design reviews, mentoring, technical roadmaps, production incidents, knowledge transfer, and presenting designs to senior leadership. Also delivered Spark MLlib training for an external company.
Currently exploring
Applying distributed-systems and data-platform principles to AI and LLM workloads. Current work is hands-on rather than deep production LLM experience.
Stack Overflow
Long-standing technical contributions on Python, Spark, PySpark, Spark SQL, DataFrames and Django.
Have a difficult distributed-systems problem?
Let’s discuss the architecture, the bottleneck or the scalability challenge. If you can describe the symptom and the shape of the system, that is usually enough to start.
Technical writing
Notes on Cassandra, Spark, Kafka and distributed systems, mostly written after working through the problem.
-
Inefficient Cassandra Query
The problem with taking over an existing project in between is there are always blindspots you aren’t aware of. This week a cassandra issue was reported on customer site which stated that a particular cassandra...
-
Flexible Scala Import
For a java developer scala is pretty much familiar - you can code without going deep into the details. Which can do the job, but magic happens when you explore it and try to write...
-
Scala List or ListBuffer
Every time I wanted to use scala collection, a question would popup - whether to use immutable collection as var or use mutable collection. As per scala collection performance they seem pretty identical, but still...
-
Kafka Consumer using Akka-Core
I faced this issue when working on a project last week. So I had to add a Kafka consumer in the project in order to write an integration test case. Now kafka consumer is pretty...