API Gateway
API Gateway
API Gateway is one of the most important components in a microservices architecture.
In a microservices system, the client should not need to know about every individual backend service. Instead, we introduce an API Gateway that acts as the entry point for client requests.
1. What is an API Gateway?
An API Gateway acts as a single entry point between the client and backend microservices.
Its primary responsibility is to:
Accept the client's API request and route it to the correct backend service based on the API endpoint.

For example:
1GET /api/invoiceThe API Gateway understands that this request belongs to the Invoice Microservice.
Why do we need an API Gateway?
Without an API Gateway, the client needs to communicate directly with each backend service such as the Order, Invoice, and Sales services. This tightly couples the client with the backend architecture. With an API Gateway, the client communicates only with the API Gateway, which handles routing requests to the appropriate backend services.

2. API Gateway vs Load Balancer
Now that we know that the API Gateway routes a request to the correct microservice, a natural question comes up:
If a Load Balancer routes traffic, then how is it different from an API Gateway?
This is one of the most common API Gateway interview questions.
The important difference is what they are routing between.
API Gateway
An API Gateway decides which microservice should receive the request.
Load Balancer
A Load Balancer generally distributes traffic between multiple instances of the same service.
Simple Difference
| API Gateway | Load Balancer |
|---|---|
| Routes requests between different services | Distributes traffic between instances |
| Understands API endpoints | Primarily distributes traffic |
/order → Order Service | Order Service → Order Instance 1/2/3 |
| Can provide authentication, rate limiting, etc. | Mainly handles traffic distribution |
So:
API Gateway decides WHICH service should handle the request.
Load Balancer decides WHICH INSTANCE of that service should handle the request.
3. API Composition
An API Gateway can route requests to different services, but in a microservices architecture, a single page may still require data from multiple services. For example, an e-commerce My Orders page might need information from the Order, Product, and Payment services.
If the client calls each service separately, it becomes responsible for coordinating multiple API calls and combining the responses. API Composition solves this by allowing the API Gateway or a dedicated service to call the required microservices, combine their responses, and return a single response to the client.
What is API Composition?
API Composition means:
The API Gateway calls multiple backend services, combines their responses, and returns a single response to the client.

Example
Suppose the client calls:
1GET /api/my-ordersThe API Gateway may internally call multiple services, such as the Product Service, Invoice Service, Review Service, and Recommendation Service, to gather all the required data. It then combines their responses and sends a single response back to the client.
It collects their responses and returns:
1{2 "product": {},3 "invoice": {},4 "reviews": [],5 "recommendations": []6}The client only makes one API call.
API Composition = Multiple backend API calls → One client response
4. Authentication
Now the API Gateway is already sitting between the client and all the microservices.
This leads to another important question:
Can we perform authentication at the API Gateway instead of implementing authentication separately in every microservice?
Yes.
The API Gateway can authenticate the client before forwarding the request to the backend services.
For example, using an OAuth 2.0-style flow:

Why authenticate at the Gateway?
Without centralized authentication, each microservice may need to handle authentication separately. This can lead to the same authentication logic being repeated across multiple services.
With an API Gateway, authentication can be handled at a central entry point. The gateway verifies the request and rejects invalid or unauthorized requests before forwarding valid requests to the microservices.
Authenticate at the entry point, then allow valid requests to continue.
5. Rate Limiting
Authentication controls who can access the API.
But another problem remains:
What if a valid client sends too many requests?
For example:
1Client2 |3 | 10,000 requests4 ▼5API Gateway6 |7 ▼8Backend ServicesThis can overload the system.
Therefore, API Gateways also provide rate limiting and throttling.
Some important mechanisms include:
- Burst limits
- API throttling
- IP-based blocking
- API queues
5.1 Burst Limit
A burst limit controls how much traffic can be handled during a sudden traffic spike.
For example:
1Burst Limit = 500During a traffic spike, the gateway can handle the configured amount of concurrent traffic.
When the limit is exceeded, requests may receive:
1HTTP 4292Too Many RequestsBurst limiting is performed at the API Gateway before the request is forwarded to the backend services. If the incoming traffic exceeds the configured burst capacity, the gateway can reject requests with HTTP 429 (Too Many Requests)
Conceptually:
1 API Gateway2 |3 +----------+----------+4 | | |5 Request Request Request6 | | |7 +----------+----------+8 |9 Burst Limit10 |11 Limit exceeded12 |13 ▼14 HTTP 4295.2 API Throttling
Burst limits handle sudden traffic spikes.
But sometimes we want a more specific rule:
How many requests is a particular user or application allowed to make?
This is where API throttling is useful.
For example:
1/api/invoice2 3Maximum:410 requests / minute / userThen:
1Request 1 ✓2Request 2 ✓3Request 3 ✓4...5Request 10 ✓6Request 11 ✗The 11th request can be blocked because the user has exceeded the allowed request rate.
Throttling can be applied at different levels:
1API2User3Application4Request Rate5.3 IP-Based Blocking
Sometimes we want to block requests from a specific IP address.
For example:
1Client2 |3 | IP: X.X.X.X4 ▼5API Gateway6 |7 | IP blocked8 ▼9Request rejectedThis can be another layer of traffic protection.
5.4 API Queues
Now consider a situation where a huge number of requests arrive at the same time. Instead of allowing all of them to reach the backend immediately, the system can limit the incoming traffic and place the allowed requests into a queue. The queue then releases requests gradually based on the capacity of the backend service.
For example, during an e-commerce flash sale, thousands of users may try to place an order at the same time. The rate limiter controls how many requests are allowed to enter the system, while the queue holds the accepted requests and sends them to the Order Service gradually based on its processing capacity. This prevents a sudden spike in traffic from overwhelming the backend.

6. Service Discovery
Now the API Gateway knows which microservice should receive the request.
But another problem appears:
Where is that microservice actually running?
In a microservices architecture, services can scale up and down.
When instances are created or removed, their IP addresses and ports can change.
Therefore, we need something that keeps track of the current location of services.
This is the job of Service Discovery.
How Service Discovery Works ?
Service Discovery maintains information about available service instances.
When a client sends a request, the API Gateway first determines which microservice should handle it. For example, if the client requests /api/order, the gateway needs to find the available instances of the Order Service.
The API Gateway queries the Service Discovery (Service Registry) to find the currently available and healthy instances of the Order Service. The registry might return multiple instances. The gateway then forwards the request to one of these healthy instances, typically through a Load Balancer, which distributes requests across the available instances.

There are two approaches discussed for maintaining this information.
Approach 1: Service Registers Itself
Whenever a microservice starts, it registers itself with Service Discovery.
Approach 2: Health Checks
Another approach is for Service Discovery to continuously perform health checks.
If a service stops responding, Service Discovery removes that instance from the active list.
Therefore, only healthy service locations remain available.

Examples of technologies used for service discovery include Eureka, Consul, ZooKeeper, etcd, and Kubernetes Service Discovery. In a Spring Cloud architecture, Eureka can act as the service registry.
7. Other API Gateway Responsibilities
So far, we have seen how an API Gateway handles routing, API composition, authentication, rate limiting, and service discovery. It can also perform several supporting tasks such as request/response transformation, caching, and centralized logging.

8. If API Gateway Is a Single Entry Point, How Does It Handle Millions of Requests?
At this point, we know that the API Gateway is the entry point.
This creates an important system-design question:
If millions of requests come through one API Gateway, won't the API Gateway itself become a bottleneck or single point of failure?
The answer is that "single entry point" does not mean there is only one physical API Gateway instance.
The API Gateway is a logical single entry point, not necessarily a single server. In a production system, it is usually deployed across multiple instances, availability zones, and sometimes regions.
-
Multiple API Gateway Instances
Instead of one gateway handling every request, multiple gateway instances run in parallel. A load balancer distributes incoming requests across these instances, allowing the system to handle much higher traffic. More instances can also be added through auto-scaling when traffic increases. -
Multiple Availability Zones
API Gateway instances can be distributed across multiple Availability Zones. If one AZ fails, traffic can be routed to instances in another AZ, improving availability and fault tolerance. -
Multiple Regions
For large-scale applications, API Gateway instances can be deployed across multiple geographic regions. Traffic can be distributed across these regions, and if one region becomes unavailable, traffic can be redirected to another region. -
Load Balancers
Load balancers distribute incoming requests across multiple API Gateway instances. This prevents a single instance from becoming overloaded and can also stop sending traffic to unhealthy instances. -
Service Discovery
After receiving a request, the API Gateway needs to know where the required microservice is running. Service discovery keeps track of available and healthy service instances so that the gateway can route requests to the correct instance. -
DNS-Based Traffic Distribution
DNS services such as AWS Route 53 or Azure Traffic Manager can distribute users across different regions or endpoints. This provides an additional layer of traffic distribution before requests reach the API Gateway.

Availability Zone
Now we have solved the problem of distributing traffic across multiple service instances. But another problem appears: what happens if the infrastructure hosting these services fails?
This is where Availability Zones (AZs) come into the picture. A Region contains multiple isolated Availability Zones, and services can be deployed across them. If one Availability Zone fails, traffic can be redirected to healthy instances in another AZ, allowing the application to continue serving users. This provides high availability and fault tolerance within a region.
Multi-AZ API Gateway Architecture
Now we can deploy our API Gateway and microservices across multiple Availability Zones (AZs) within the same region.
Each AZ runs its own API Gateway and service instances. Traffic is distributed across these AZs, so the application does not depend on a single AZ.
If AZ1 fails, traffic can be redirected to AZ2, allowing the application to continue serving requests.
This improves the system's availability and fault tolerance.
Multi-Region Architecture
Multi-AZ protects the application from failures within a region. But if an entire region becomes unavailable, we need another level of redundancy. This is where Multi-Region Architecture is used.
The application can be deployed across multiple regions, with each region having its own API Gateway, Load Balancers, Service Discovery, and microservice instances. Traffic can be distributed between regions using global DNS or traffic-routing services.
If one Availability Zone fails, another AZ within the same region can continue serving traffic. If the entire region fails, traffic can be redirected to another healthy region, allowing the application to continue serving users.
15. DNS-Based Traffic Distribution
Now we have multiple regions.
This creates another question:
If there are multiple API Gateways in different regions, who decides which region receives the client's request?
Now that we have multiple regions, another question arises: how does the system decide which region should handle a client’s request? This is where DNS-based traffic distribution comes into play. Services such as AWS Route 53 and Azure Traffic Manager can direct users to different regions based on configured routing rules.
For example, the traffic manager can consider factors such as latency, geographic location, or compliance requirements when selecting a region. Once a region is selected, the request is sent to the API Gateway in that region. This allows traffic to be distributed across regions and helps the application remain available if a region becomes unavailable.
Complete API Gateway Architecture
Now we can combine everything we have learned.

Remember the responsibilities
| Component | Main Responsibility |
|---|---|
| DNS / Traffic Manager | Choose the appropriate region |
| API Gateway | Route request to the correct service |
| Authentication | Verify the client/token |
| Rate Limiting | Control request rate |
| Service Discovery | Find current service locations |
| Load Balancer | Distribute traffic across service instances |
| Microservice | Process the actual business request |
One-Line Mental Model
DNS chooses the region → API Gateway chooses the service → Service Discovery finds its location → Load Balancer chooses the instance → Microservice processes the request.
Interview Answer: What is an API Gateway?
An API Gateway is a single entry point between clients and backend microservices in a microservices architecture.
Instead of the client directly communicating with multiple services, it sends requests to the API Gateway. The gateway identifies which microservice should handle the request and routes it accordingly. For example, a request to /api/order would be routed to the Order Service.
Apart from routing, an API Gateway can also handle common responsibilities such as authentication, rate limiting, API composition, service discovery, caching, and logging.
In a production system, multiple API Gateway instances can run across Availability Zones and regions, so the gateway itself does not become a bottleneck or single point of failure.
Simple Example
Suppose an e-commerce application has:
- Order Service
- Payment Service
- Product Service