Principal Engineer & Independent Technical Consultant

Ayush Vatsyayan

Distributed Systems · Event-Driven Architecture · Data Platforms · Performance & Scalability

I help engineering teams design, troubleshoot and scale distributed systems, event-driven platforms and data-intensive applications.

Consulting Case Studies Technical Writing GitHub Contact

Problems I help solve

Most of the engagements I take on start as a symptom somewhere in the stack and a bottleneck somewhere else. The work is tracing the path between them.

Independent architecture review

A second opinion on an existing or proposed architecture, before or after a significant investment.

Modernize Legacy Technology

Modernize legacy technology and identify practical analytics opportunities acros engineering and operations.

Difficult production problems

A live incident or recurring failure where the real bottleneck has not been identified yet.

Slow or unreliable distributed systems

Throughput is not keeping up, scaling is unpredictable, or behaviour under load is hard to reason about.

Kafka consumer lag and event pipelines

Lag keeps growing, processing throughput has plateaued, or partitioning and topic configuration need a second look.

Cassandra performance and data modelling

Queries degrade at scale and read performance are hurting, or a data model needs reviewing.

Selected case studies

Sanitized accounts of real engagements, focused on the diagnostic process and the architectural reasoning rather than company specifics or invented metrics.

  • A commit pipeline that was too slow

    A single-partitioned Kafka topic, synchronous blocking in the application thread and a cache read/write/delete pattern in Cassandra all contributed to the same problem.

    Scala · Akka/Pekko · Kafka · Cassandra · Docker

  • Kafka lag traced back to Cassandra tombstones

    Rising consumer lag turned out to be a data lifecycle problem. Expiring rows individually produced enough tombstones to degrade reads feeding the pipeline.

    Scala · Kafka · Cassandra · Data modelling

  • A reporting system that failed at scale

    Fixing it required changing the data-access model and the user workflow, not optimising the existing queries.

    Java · MySQL · BIRT · Reporting architecture

Read the case studies

Engagement model

I am currently running a focused review session as a first, low-commitment way to work together. Have a distributed-system performance or scalability problem you cannot pin down?

Distributed Systems Architecture & Technical Review

  • 60–90 minute technical session
  • Review of the architecture, the problem and available evidence
  • Discussion of relevant logs, code and design
  • Identification of the likely bottleneck and its contributing factors
  • Written findings
  • Prioritized recommendations and follow-up Q&A

Initial rate: ₹5,000

This is only a diagnosis and review engagement.

Consulting details

Background

18 years of production systems

Senior and principal engineering across telecom, analytics, machine learning, big data and enterprise platforms, mostly as an individual contributor.

Event-driven platforms at scale

Designed and led an event-driven ETL platform with high-throughput aggregation pipelines, and distributed event-streaming engines modelling complex Layer-3 network topologies. Scala, Akka/Pekko, Kafka, Cassandra, REST APIs, Docker, Kubernetes.

Data platforms and ML

Moved a text analytics platform from Excel and small local datasets onto Spark, PySpark, Hadoop and Scala. Earlier work across Hadoop, Hive, Spark, Cassandra and telecom analytics at Nokia and Alcatel-Lucent.

Working with engineering teams

Architecture and design reviews, mentoring, technical roadmaps, production incidents, knowledge transfer, and presenting designs to senior leadership. Also delivered Spark MLlib training for an external company.

Currently exploring

Applying distributed-systems and data-platform principles to AI and LLM workloads. Current work is hands-on rather than deep production LLM experience.

Stack Overflow

Long-standing technical contributions on Python, Spark, PySpark, Spark SQL, DataFrames and Django.

stackoverflow.com/users/6065591/ayush-vatsyayan

About me

Have a difficult distributed-systems problem?

Let’s discuss the architecture, the bottleneck or the scalability challenge. If you can describe the symptom and the shape of the system, that is usually enough to start.

vatsyayan.ayush@gmail.com Contact details LinkedIn

Technical writing

Notes on Cassandra, Spark, Kafka and distributed systems, mostly written after working through the problem.

  • Inefficient Cassandra Query

    The problem with taking over an existing project in between is there are always blindspots you aren’t aware of. This week a cassandra issue was reported on customer site which stated that a particular cassandra...

  • Flexible Scala Import

    For a java developer scala is pretty much familiar - you can code without going deep into the details. Which can do the job, but magic happens when you explore it and try to write...

  • Scala List or ListBuffer

    Every time I wanted to use scala collection, a question would popup - whether to use immutable collection as var or use mutable collection. As per scala collection performance they seem pretty identical, but still...

  • Kafka Consumer using Akka-Core

    I faced this issue when working on a project last week. So I had to add a Kafka consumer in the project in order to write an integration test case. Now kafka consumer is pretty...

All writing