Chaturmind
LearnDSASystem DesignInterview PrepDevOpsEngineering GrowthBlog
Start learning
Chaturmind

Structured learning paths for engineers who want to go deep. Written by practitioners.

Learn

  • Java
  • DSA
  • System Design
  • Spring Boot
  • AI / ML
  • DevOps
  • Engineering Growth
  • Java Interview Prep

Company

  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 Chaturmind. All rights reserved.

Built for engineers who want to go deep.


← Java Interview Prep: 8+ Years (Senior & Lead)

Expert Core Java

  • Tricky Java Output, Operators & OOP Edge Cases — Interview Questions
  • Tricky Exceptions, Memory & Keyword Questions — Interview Questions
  • Classic Java Language Questions, Senior-Grade Answers — Interview Questions
  • Classic Collections, Threads & JDK APIs, Senior-Grade Answers — Interview Questions
  • Reflection, Dynamic Proxies, final & Modern OOP Design — Interview Questions

JVM Internals & Performance

  • Class Loading, Bytecode & Object Layout — Interview Questions
  • JIT Compilation & Runtime Optimisations — Interview Questions
  • Garbage Collectors Deep Dive — Interview Questions
  • JVM Tuning, GC Logs & Memory Footprint — Interview Questions
  • Memory Leaks, OutOfMemoryErrors & Profiling Tools — Interview Questions
  • Modules, Agents & Advanced JVM APIs — Interview Questions

Collections & Concurrency at Scale

  • Collections Internals & Complexity — Interview Questions
  • Iterators, Comparators & Ordering Contracts — Interview Questions
  • Concurrent Collections, Queues & Lock-Free Structures — Interview Questions
  • Threads, Executors & ForkJoin Internals — Interview Questions
  • Locks, Atomics, CAS & Synchronizers — Interview Questions
  • Java Memory Model, volatile, Fences & ThreadLocal — Interview Questions
  • Deadlock, Livelock, Starvation & Concurrent Design — Interview Questions
  • CompletableFuture, Parallel Streams & Non-Blocking I/O — Interview Questions

Modern Java (8 to 21+)

  • Lambdas & Functional Interfaces Internals — Interview Questions
  • Streams & Collectors Deep Dive — Interview Questions
  • Optional & Interface Default/Static Methods — Interview Questions
  • Java 9–25 Features & Virtual Threads — Interview Questions

Design Patterns, SOLID & Clean Code

  • Design Pattern Trade-offs & Combinations — Interview Questions
  • SOLID, Clean Code & Anti-Patterns — Interview Questions

Spring & Spring Boot Internals

  • IoC, Dependency Injection & Bean Lifecycle Internals — Interview Questions
  • Spring AOP, Proxies & @Async Internals — Interview Questions
  • Spring Configuration, Auto-Configuration & Custom Starters — Interview Questions
  • Spring MVC & REST Internals, Exception Frameworks — Interview Questions
  • Spring Security Advanced Internals — Interview Questions
  • Spring WebFlux, Reactor & R2DBC — Interview Questions
  • Spring Cloud, Observability & Distributed Tracing — Interview Questions
  • Spring Boot 3, Native Images & Production Scenarios — Interview Questions

JPA, Hibernate & Databases at Scale

  • Spring Data JPA — Queries, Projections, Custom Repositories & Locking — Interview Questions
  • JPA Entity Mapping, Associations & Cascades — Interview Questions
  • JPQL vs Native Queries in Depth — Interview Questions
  • Hibernate Caching — First-Level, Second-Level & Query Cache — Interview Questions
  • Lazy vs Eager Loading, LazyInitializationException & N+1 — Interview Questions
  • JPA Transactions, Propagation, Isolation & Dirty Checking — Interview Questions
  • SQL vs NoSQL, Indexing & Query Tuning — Interview Questions
  • Database Scaling, Replication, Pooling & Consistency Models — Interview Questions
  • Redis, Search, Time-Series, CDC & Transactional Data Modelling — Interview Questions

Testing Strategy & API Design

  • Spring Boot Test Slices, Context & Test Strategy — Interview Questions
  • Testing Web, Persistence, Security, Async & Messaging in Spring Boot — Interview Questions
  • JUnit 5 & Mockito, Advanced — Interview Questions
  • MockMvc, WebTestClient & Testcontainers in Depth — Interview Questions
  • REST Principles, Status Codes & Resource Design — Interview Questions
  • OpenAPI, Validation Errors, API Versioning & GraphQL — Interview Questions

Build, DevOps & Cloud

  • Maven & Gradle at Scale — Interview Questions
  • Git, CI/CD Pipelines & Release Safety — Interview Questions
  • Docker & Kubernetes for Java Engineers — Interview Questions
  • Quality Gates, Artifact Repositories & Secrets Management — Interview Questions
  • AWS Deployment & Scaling for Spring Boot — Interview Questions
  • Multi-Cloud Deployment, High Availability, Cost & Cloud Troubleshooting — Interview Questions

Kafka & Messaging

  • Kafka Internals & Delivery Semantics — Interview Questions
  • Spring Kafka — Error Handling, DLQs, Schemas & Operations — Interview Questions
  • RabbitMQ, JMS & Messaging Models — Interview Questions

Microservices & Architecture

  • Distributed Systems Fundamentals — CAP, Consistency, Availability & SLOs — Interview Questions
  • DDD, Hexagonal Architecture & Service Boundaries — Interview Questions
  • Event-Driven Architecture, CQRS, Event Sourcing, Sharding & Idempotency — Interview Questions
  • Rate Limiting, Resilience, Caching at Scale & Chaos Engineering — Interview Questions
  • Files, Documents & Internationalisation in Java Backends — Interview Questions
  • WebSockets, Schedulers, Notifications & Real-Time Pipelines — Interview Questions

System Design Scenarios

  • Booking Systems, CRS, Inventory & Concurrency Control — Interview Questions
  • Dynamic Pricing & Rule Engines — Interview Questions
  • Partner Integrations — OTA Sync, Retries, Webhooks, Reconciliation & Bulk Data — Interview Questions
  • Designing Caches & Rate Limiters — Interview Questions
  • Event-Driven Architecture, Kafka at Scale, IoT & Real-Time Pipelines — Interview Questions
  • Observability, Logging, Alerting & Audit Systems — Interview Questions
  • Multi-Tenant SaaS, Identity & Platform Services — Interview Questions
  • Search, Notifications, Chat, Fraud Detection & Workflows — Interview Questions
  • Extreme Scale, 99.99% Availability, DR & Project Deep-Dive Stories — Interview Questions

Security for Senior Engineers

  • Tokens, OAuth2 PKCE, Web Attacks & API Security — Interview Questions
  • TLS, mTLS, Zero Trust, Secrets, DDoS & Privacy Compliance — Interview Questions

Leadership & Behavioural

  • Leadership Style, Motivation & Team Health — Interview Questions
  • Delivery, Planning & Decisions Under Uncertainty — Interview Questions
  • Problem Solving, Growth & Career Stories — Interview Questions
  • Stakeholder Communication, Ethics & Compliance — Interview Questions
  • Mentoring, Knowledge Sharing & Code Reviews — Interview Questions
  • Agile & Scrum Practices for Senior Engineers — Interview Questions
  • Architecture Decision-Making — Interview Questions
  • Conflict Resolution & Difficult Conversations — Interview Questions
Chaturmind
← Java Interview Prep: 8+ Years (Senior & Lead)

Expert Core Java

  • Tricky Java Output, Operators & OOP Edge Cases — Interview Questions
  • Tricky Exceptions, Memory & Keyword Questions — Interview Questions
  • Classic Java Language Questions, Senior-Grade Answers — Interview Questions
  • Classic Collections, Threads & JDK APIs, Senior-Grade Answers — Interview Questions
  • Reflection, Dynamic Proxies, final & Modern OOP Design — Interview Questions

JVM Internals & Performance

  • Class Loading, Bytecode & Object Layout — Interview Questions
  • JIT Compilation & Runtime Optimisations — Interview Questions
  • Garbage Collectors Deep Dive — Interview Questions
  • JVM Tuning, GC Logs & Memory Footprint — Interview Questions
  • Memory Leaks, OutOfMemoryErrors & Profiling Tools — Interview Questions
  • Modules, Agents & Advanced JVM APIs — Interview Questions

Collections & Concurrency at Scale

  • Collections Internals & Complexity — Interview Questions
  • Iterators, Comparators & Ordering Contracts — Interview Questions
  • Concurrent Collections, Queues & Lock-Free Structures — Interview Questions
  • Threads, Executors & ForkJoin Internals — Interview Questions
  • Locks, Atomics, CAS & Synchronizers — Interview Questions
  • Java Memory Model, volatile, Fences & ThreadLocal — Interview Questions
  • Deadlock, Livelock, Starvation & Concurrent Design — Interview Questions
  • CompletableFuture, Parallel Streams & Non-Blocking I/O — Interview Questions

Modern Java (8 to 21+)

  • Lambdas & Functional Interfaces Internals — Interview Questions
  • Streams & Collectors Deep Dive — Interview Questions
  • Optional & Interface Default/Static Methods — Interview Questions
  • Java 9–25 Features & Virtual Threads — Interview Questions

Design Patterns, SOLID & Clean Code

  • Design Pattern Trade-offs & Combinations — Interview Questions
  • SOLID, Clean Code & Anti-Patterns — Interview Questions

Spring & Spring Boot Internals

  • IoC, Dependency Injection & Bean Lifecycle Internals — Interview Questions
  • Spring AOP, Proxies & @Async Internals — Interview Questions
  • Spring Configuration, Auto-Configuration & Custom Starters — Interview Questions
  • Spring MVC & REST Internals, Exception Frameworks — Interview Questions
  • Spring Security Advanced Internals — Interview Questions
  • Spring WebFlux, Reactor & R2DBC — Interview Questions
  • Spring Cloud, Observability & Distributed Tracing — Interview Questions
  • Spring Boot 3, Native Images & Production Scenarios — Interview Questions

JPA, Hibernate & Databases at Scale

  • Spring Data JPA — Queries, Projections, Custom Repositories & Locking — Interview Questions
  • JPA Entity Mapping, Associations & Cascades — Interview Questions
  • JPQL vs Native Queries in Depth — Interview Questions
  • Hibernate Caching — First-Level, Second-Level & Query Cache — Interview Questions
  • Lazy vs Eager Loading, LazyInitializationException & N+1 — Interview Questions
  • JPA Transactions, Propagation, Isolation & Dirty Checking — Interview Questions
  • SQL vs NoSQL, Indexing & Query Tuning — Interview Questions
  • Database Scaling, Replication, Pooling & Consistency Models — Interview Questions
  • Redis, Search, Time-Series, CDC & Transactional Data Modelling — Interview Questions

Testing Strategy & API Design

  • Spring Boot Test Slices, Context & Test Strategy — Interview Questions
  • Testing Web, Persistence, Security, Async & Messaging in Spring Boot — Interview Questions
  • JUnit 5 & Mockito, Advanced — Interview Questions
  • MockMvc, WebTestClient & Testcontainers in Depth — Interview Questions
  • REST Principles, Status Codes & Resource Design — Interview Questions
  • OpenAPI, Validation Errors, API Versioning & GraphQL — Interview Questions

Build, DevOps & Cloud

  • Maven & Gradle at Scale — Interview Questions
  • Git, CI/CD Pipelines & Release Safety — Interview Questions
  • Docker & Kubernetes for Java Engineers — Interview Questions
  • Quality Gates, Artifact Repositories & Secrets Management — Interview Questions
  • AWS Deployment & Scaling for Spring Boot — Interview Questions
  • Multi-Cloud Deployment, High Availability, Cost & Cloud Troubleshooting — Interview Questions

Kafka & Messaging

  • Kafka Internals & Delivery Semantics — Interview Questions
  • Spring Kafka — Error Handling, DLQs, Schemas & Operations — Interview Questions
  • RabbitMQ, JMS & Messaging Models — Interview Questions

Microservices & Architecture

  • Distributed Systems Fundamentals — CAP, Consistency, Availability & SLOs — Interview Questions
  • DDD, Hexagonal Architecture & Service Boundaries — Interview Questions
  • Event-Driven Architecture, CQRS, Event Sourcing, Sharding & Idempotency — Interview Questions
  • Rate Limiting, Resilience, Caching at Scale & Chaos Engineering — Interview Questions
  • Files, Documents & Internationalisation in Java Backends — Interview Questions
  • WebSockets, Schedulers, Notifications & Real-Time Pipelines — Interview Questions

System Design Scenarios

  • Booking Systems, CRS, Inventory & Concurrency Control — Interview Questions
  • Dynamic Pricing & Rule Engines — Interview Questions
  • Partner Integrations — OTA Sync, Retries, Webhooks, Reconciliation & Bulk Data — Interview Questions
  • Designing Caches & Rate Limiters — Interview Questions
  • Event-Driven Architecture, Kafka at Scale, IoT & Real-Time Pipelines — Interview Questions
  • Observability, Logging, Alerting & Audit Systems — Interview Questions
  • Multi-Tenant SaaS, Identity & Platform Services — Interview Questions
  • Search, Notifications, Chat, Fraud Detection & Workflows — Interview Questions
  • Extreme Scale, 99.99% Availability, DR & Project Deep-Dive Stories — Interview Questions

Security for Senior Engineers

  • Tokens, OAuth2 PKCE, Web Attacks & API Security — Interview Questions
  • TLS, mTLS, Zero Trust, Secrets, DDoS & Privacy Compliance — Interview Questions

Leadership & Behavioural

  • Leadership Style, Motivation & Team Health — Interview Questions
  • Delivery, Planning & Decisions Under Uncertainty — Interview Questions
  • Problem Solving, Growth & Career Stories — Interview Questions
  • Stakeholder Communication, Ethics & Compliance — Interview Questions
  • Mentoring, Knowledge Sharing & Code Reviews — Interview Questions
  • Agile & Scrum Practices for Senior Engineers — Interview Questions
  • Architecture Decision-Making — Interview Questions
  • Conflict Resolution & Difficult Conversations — Interview Questions
HomeLearnJava Interview PrepJava Interview Prep: 8+ Years (Senior & Lead)Microservices & Architecture
✓ FreeAdvanced· 9 min read

Distributed Systems Fundamentals — CAP, Consistency, Availability & SLOs — Interview Questions

Horizontal vs vertical scaling, the CAP theorem stated correctly (and PACELC), a real-world consistency-vs-availability decision, consistency models, designing for high availability and fault tolerance, what 99.99% really requires, SLI/SLO/SLA and error budgets, two-phase commit vs sagas, and distributed locking.

Published September 25, 2026


How to use this lesson

These are the conceptual foundations interviewers use to judge whether your designs are principled. State the theory precisely (CAP is often misquoted), then connect it to a concrete decision you'd make, and its cost.

Q1. What's the difference between horizontal and vertical scaling?

Short answer:

  • Vertical scaling (scale up): a bigger machine (more CPU, RAM, faster disks). It's simple (no code changes, no distribution problems), but it has hard ceilings, cost that grows non-linearly, a single point of failure, and often downtime to resize.
  • Horizontal scaling (scale out): more machines or instances, behind load balancing or partitioning. It's near-linear capacity growth, and gives fault tolerance (losing one node isn't fatal), and elasticity (autoscaling). It requires: statelessness, or partitioned state; load balancing; and handling distributed-systems problems (consistency, coordination, network failures).

In practice: scale the stateless tier horizontally. Scale databases vertically first (it's often cheaper and simpler), then with read replicas, partitioning and sharding when you must.

Learn it in depth → Horizontal vs Vertical Scaling

Q2. What is the CAP theorem, and how does it affect distributed design?

Short answer: For a distributed data store, when a network partition (P) happens, you must choose between:

  • Consistency (C): every read sees the latest write (linearisability);
  • Availability (A): every request to a non-failed node gets a (non-error) response.

Partitions aren't optional in real networks, so the practical choice is CP vs AP during a partition:

  • CP systems (etcd, ZooKeeper, HBase, Spanner, and a relational primary with synchronous replication) refuse or block some requests, to stay consistent;
  • AP systems (Cassandra and DynamoDB with eventual consistency, DNS, shopping-cart-style designs) keep answering, possibly with stale or conflicting data, and reconcile later.

Common trap: the popular phrasing "pick any two of three" is misleading. You can't give up P in a distributed system. PACELC refines it: if Partitioned, choose A or C; Else (normal operation), choose Latency or Consistency. That explains why many systems trade consistency for latency even with no partition.

The design impact: choose per use case, not per system. Strong consistency for money, inventory and uniqueness constraints. Availability and eventual consistency for catalogues, feeds, recommendations and analytics.

Learn it in depth → CAP Theorem

Q3. Give a real-world scenario where you had to choose between consistency and availability.

Short answer: Use the STAR structure. A strong example: during festival sales on an e-commerce platform:

  • Product browsing, search and prices shown are served from caches, replicas and the search index. We accepted eventual consistency (a price or stock badge could be seconds stale), so the site stayed available and fast, even when the primary database or a region was under stress. That's AP.
  • Checkout and inventory reservation stayed strongly consistent (CP): an atomic conditional decrement on the primary, so there was no overselling. When the inventory database was unreachable, checkout failed fast with a friendly retry message, rather than accepting orders we couldn't fulfil.
  • The result: high availability for 95% of the traffic, correct orders, and clear customer messaging. We monitored oversell count (0), checkout error rate, and staleness.

Key points to cover:

  • It's rarely "the whole system is AP". Different operations get different guarantees.

Q4. What consistency models exist in distributed systems?

Short answer: From strongest to weakest:

  • Linearisability (strong consistency): operations appear instantaneous at some point between their invocation and completion, in real-time order. It needs consensus (Raft or Paxos) or a single leader.
  • Sequential consistency: everyone sees the same order of operations, respecting each client's program order, but not necessarily real time.
  • Causal consistency: causally related operations (a reply after the post) are seen in order by everyone. Concurrent operations may differ.
  • Session guarantees (practical, client-centric): read-your-writes, monotonic reads, monotonic writes, writes-follow-reads.
  • Bounded staleness: reads lag by at most k versions or t seconds (Cosmos DB offers this).
  • Eventual consistency: if updates stop, the replicas converge. There are no ordering guarantees in the meantime.

Database transaction isolation levels (read committed, repeatable read, serialisable) are a different axis: they concern concurrent transactions, not replica convergence. Strict serialisability combines serialisability with linearisability (Spanner).

Q5. How do you design a system for high availability and fault tolerance? What does 99.99% availability require?

Short answer:

  • The principles:
    • eliminate single points of failure (redundancy at every layer: instances, zones, load balancers, databases with automatic failover, and brokers with replication);
    • isolate failures (bulkheads, circuit breakers, timeouts, cell or shard architectures);
    • degrade gracefully (fallbacks, feature shedding);
    • detect and recover automatically (health checks, self-healing orchestration, automated failover);
    • deploy safely (canaries, fast rollback);
    • and test failure (chaos engineering, game days).
  • 99.99% ("four nines") allows about 52.6 minutes of downtime per year, and about 4.3 minutes per month. That means:
    • no manual steps in the recovery path: failover must be automatic, and detected within seconds;
    • multi-AZ everything (and often multi-region for the critical paths), with every dependency at least as available (serial dependencies multiply: 5 dependencies at 99.99% give about 99.95% overall), which pushes you toward async decoupling and redundancy;
    • zero-downtime deployments and migrations, with rapid rollback;
    • capacity headroom (N+1 or N+2), and load shedding;
    • SLO-based alerting and on-call practices, with fast incident response.
  • Cost: each extra nine costs a lot more. Agree on the target with the business, per user journey.

Q6. What are SLIs, SLOs and SLAs? What's an error budget?

Short answer:

  • SLI (Service Level Indicator): a measured metric of user-perceived reliability. For example, the proportion of successful requests, or the proportion of requests faster than 300 ms, measured at the load balancer or the client.

  • SLO (Service Level Objective): the internal target for an SLI over a window. For example, "99.9% of checkout requests succeed over 28 days", or "99% under 300 ms". It drives engineering priorities.

  • SLA (Service Level Agreement): a contractual promise to customers, with consequences (credits or penalties). It's usually looser than the SLO, so you have a buffer.

  • Error budget: 1 − SLO, the allowed unreliability. At 99.9% over 28 days, that's about 40 minutes of failures. How it's used:

    • while there's budget left, ship features faster (take risks);
    • when it's exhausted (or burning fast), freeze risky releases, and prioritise reliability work;
    • alert on the burn rate (multi-window, multi-burn-rate alerts), not on raw thresholds.

    It aligns product and operations around one number.

Learn it in depth → Alerting Strategy

Q7. Two-phase commit vs sagas?

Short answer:

  • Two-phase commit (2PC/XA):

    1. the coordinator asks every participant to prepare (vote yes or no, holding locks);
    2. if all vote yes, it sends commit, otherwise abort.

    It gives atomicity and strong consistency across resources. The drawbacks:

    • it's blocking: if the coordinator fails after the prepare phase, participants hold locks indefinitely (in-doubt transactions);
    • latency from several round trips;
    • it reduces availability (every participant must be up);
    • it's poorly supported by modern infrastructure (Kafka, many NoSQL and cloud databases);
    • it creates tight coupling.

    It fits within one trust boundary, with XA-capable resources (two databases in one application server), rarely across microservices.

  • Sagas: a sequence of local transactions, with compensating actions, coordinated by events (choreography) or an orchestrator. Non-blocking, highly available, loosely coupled, and it works with any technology. The costs:

    • no isolation (intermediate states are visible, so use semantic locks or pending states);
    • eventual consistency;
    • compensations must be designed and idempotent;
    • more complex reasoning and testing.
  • In microservices, use sagas (plus the outbox and idempotency). Keep invariants that need atomicity inside a single service's database.

Learn it in depth → Two-Phase Commit

Q8. What is distributed locking, and when is it (un)safe?

Short answer: A mutual-exclusion lock shared by processes on different machines, used to ensure that only one instance performs an action at a time (a scheduled job, a resource refresh, a leader-only task). Implementations:

  • database locks: SELECT … FOR UPDATE, advisory locks (pg_try_advisory_lock), or a lock table (ShedLock uses this for @Scheduled jobs);
  • Redis: SET key token NX PX ttl, with a token-checked release (Lua), and Redisson's RLock (a watchdog renewal);
  • consensus systems: etcd leases, ZooKeeper ephemeral sequential nodes, Consul sessions. These give stronger guarantees;
  • Kubernetes Lease objects, for leader election.

The pitfalls:

  • Leases can expire while the holder is paused (GC pauses, network delays, VM stalls), so two processes believe they hold the lock. For correctness-critical operations, use fencing tokens: a monotonically increasing number issued with each lock grant, checked by the protected resource, which rejects older tokens.
  • The Redlock debates (Kleppmann vs antirez): don't rely on a timing-based lock for safety.
  • Prefer designs that don't need locks: idempotent operations, single-writer partitioning (Kafka partition ownership), optimistic concurrency, or database constraints.

Follow-up questions this topic invites — and their answers

Q: Is a single relational database CP or AP? A: A single node isn't a distributed system in the CAP sense. With replicas: synchronous replication with failover that refuses writes during partitions is CP-ish, and asynchronous replicas serving reads are AP-ish for those reads (stale data). It depends on how you configure and use it.

Q: What's the difference between availability and durability? A: Availability means the system responds to requests now. Durability means committed data survives failures. S3, for example, offers 11 nines of durability, but lower availability SLAs. You can lose availability temporarily without losing data, and vice versa.

Q: How do you compute a composite SLO for dependent services? A: For serial hard dependencies, multiply the availabilities (for example 0.999 × 0.999 ≈ 0.998). Redundant parallel paths improve it: 1 − (1 − a)ⁿ. That's why critical paths should minimise synchronous hard dependencies.

Q: What is a "cell-based architecture"? A: Splitting the whole stack into independent cells (each serving a subset of customers), so failures and bad deployments affect only one cell's blast radius. It's used by large SaaS providers and AWS services.

Previous

RabbitMQ, JMS & Messaging Models — Interview Questions

Next

DDD, Hexagonal Architecture & Service Boundaries — Interview Questions

AI Tutor

Lesson: Distributed Systems Fundamentals — CAP, Consistency, Availability & SLOs — Interview Questions

Quick actions

AI responses can be inaccurate. Verify critical information.