API & Security

API Gateway

API Gateway

API Gateway is one of the most important components in a microservices architecture.

In a microservices system, the client should not need to know about every individual backend service. Instead, we introduce an API Gateway that acts as the entry point for client requests.

1. What is an API Gateway?

An API Gateway acts as a single entry point between the client and backend microservices.

Its primary responsibility is to:

Accept the client's API request and route it to the correct backend service based on the API endpoint.

For example:

text
1GET /api/invoice

The API Gateway understands that this request belongs to the Invoice Microservice.

Why do we need an API Gateway?

Without an API Gateway, the client needs to communicate directly with each backend service such as the Order, Invoice, and Sales services. This tightly couples the client with the backend architecture. With an API Gateway, the client communicates only with the API Gateway, which handles routing requests to the appropriate backend services.

2. API Gateway vs Load Balancer

Now that we know that the API Gateway routes a request to the correct microservice, a natural question comes up:

If a Load Balancer routes traffic, then how is it different from an API Gateway?

This is one of the most common API Gateway interview questions.

The important difference is what they are routing between.

API Gateway

An API Gateway decides which microservice should receive the request.

Load Balancer

A Load Balancer generally distributes traffic between multiple instances of the same service.

Simple Difference

API GatewayLoad Balancer
Routes requests between different servicesDistributes traffic between instances
Understands API endpointsPrimarily distributes traffic
/order → Order ServiceOrder Service → Order Instance 1/2/3
Can provide authentication, rate limiting, etc.Mainly handles traffic distribution

So:

API Gateway decides WHICH service should handle the request.

Load Balancer decides WHICH INSTANCE of that service should handle the request.

3. API Composition

An API Gateway can route requests to different services, but in a microservices architecture, a single page may still require data from multiple services. For example, an e-commerce My Orders page might need information from the Order, Product, and Payment services.

If the client calls each service separately, it becomes responsible for coordinating multiple API calls and combining the responses. API Composition solves this by allowing the API Gateway or a dedicated service to call the required microservices, combine their responses, and return a single response to the client.

What is API Composition?

API Composition means:

The API Gateway calls multiple backend services, combines their responses, and returns a single response to the client.

Example

Suppose the client calls:

text
1GET /api/my-orders

The API Gateway may internally call multiple services, such as the Product Service, Invoice Service, Review Service, and Recommendation Service, to gather all the required data. It then combines their responses and sends a single response back to the client.

It collects their responses and returns:

json
1{
2 "product": {},
3 "invoice": {},
4 "reviews": [],
5 "recommendations": []
6}

The client only makes one API call.

API Composition = Multiple backend API calls → One client response

4. Authentication

Now the API Gateway is already sitting between the client and all the microservices.

This leads to another important question:

Can we perform authentication at the API Gateway instead of implementing authentication separately in every microservice?

Yes.

The API Gateway can authenticate the client before forwarding the request to the backend services.

For example, using an OAuth 2.0-style flow:

Why authenticate at the Gateway?

Without centralized authentication, each microservice may need to handle authentication separately. This can lead to the same authentication logic being repeated across multiple services.

With an API Gateway, authentication can be handled at a central entry point. The gateway verifies the request and rejects invalid or unauthorized requests before forwarding valid requests to the microservices.

Authenticate at the entry point, then allow valid requests to continue.

5. Rate Limiting

Authentication controls who can access the API.

But another problem remains:

What if a valid client sends too many requests?

For example:

text
1Client
2 |
3 | 10,000 requests
4 ▼
5API Gateway
6 |
7 ▼
8Backend Services

This can overload the system.

Therefore, API Gateways also provide rate limiting and throttling.

Some important mechanisms include:

  • Burst limits
  • API throttling
  • IP-based blocking
  • API queues

5.1 Burst Limit

A burst limit controls how much traffic can be handled during a sudden traffic spike.

For example:

text
1Burst Limit = 500

During a traffic spike, the gateway can handle the configured amount of concurrent traffic.

When the limit is exceeded, requests may receive:

text
1HTTP 429
2Too Many Requests

Burst limiting is performed at the API Gateway before the request is forwarded to the backend services. If the incoming traffic exceeds the configured burst capacity, the gateway can reject requests with HTTP 429 (Too Many Requests)

Conceptually:

text
1 API Gateway
2 |
3 +----------+----------+
4 | | |
5 Request Request Request
6 | | |
7 +----------+----------+
8 |
9 Burst Limit
10 |
11 Limit exceeded
12 |
13 ▼
14 HTTP 429

5.2 API Throttling

Burst limits handle sudden traffic spikes.

But sometimes we want a more specific rule:

How many requests is a particular user or application allowed to make?

This is where API throttling is useful.

For example:

text
1/api/invoice
2
3Maximum:
410 requests / minute / user

Then:

text
1Request 1 ✓
2Request 2 ✓
3Request 3 ✓
4...
5Request 10 ✓
6Request 11 ✗

The 11th request can be blocked because the user has exceeded the allowed request rate.

Throttling can be applied at different levels:

text
1API
2User
3Application
4Request Rate

5.3 IP-Based Blocking

Sometimes we want to block requests from a specific IP address.

For example:

text
1Client
2 |
3 | IP: X.X.X.X
4 ▼
5API Gateway
6 |
7 | IP blocked
8 ▼
9Request rejected

This can be another layer of traffic protection.


5.4 API Queues

Now consider a situation where a huge number of requests arrive at the same time. Instead of allowing all of them to reach the backend immediately, the system can limit the incoming traffic and place the allowed requests into a queue. The queue then releases requests gradually based on the capacity of the backend service.

For example, during an e-commerce flash sale, thousands of users may try to place an order at the same time. The rate limiter controls how many requests are allowed to enter the system, while the queue holds the accepted requests and sends them to the Order Service gradually based on its processing capacity. This prevents a sudden spike in traffic from overwhelming the backend.

6. Service Discovery

Now the API Gateway knows which microservice should receive the request.

But another problem appears:

Where is that microservice actually running?

In a microservices architecture, services can scale up and down.

When instances are created or removed, their IP addresses and ports can change.

Therefore, we need something that keeps track of the current location of services.

This is the job of Service Discovery.

How Service Discovery Works ?

Service Discovery maintains information about available service instances.

When a client sends a request, the API Gateway first determines which microservice should handle it. For example, if the client requests /api/order, the gateway needs to find the available instances of the Order Service.

The API Gateway queries the Service Discovery (Service Registry) to find the currently available and healthy instances of the Order Service. The registry might return multiple instances. The gateway then forwards the request to one of these healthy instances, typically through a Load Balancer, which distributes requests across the available instances.

There are two approaches discussed for maintaining this information.

Approach 1: Service Registers Itself

Whenever a microservice starts, it registers itself with Service Discovery.

Approach 2: Health Checks

Another approach is for Service Discovery to continuously perform health checks.

If a service stops responding, Service Discovery removes that instance from the active list.

Therefore, only healthy service locations remain available.

Examples of technologies used for service discovery include Eureka, Consul, ZooKeeper, etcd, and Kubernetes Service Discovery. In a Spring Cloud architecture, Eureka can act as the service registry.

7. Other API Gateway Responsibilities

So far, we have seen how an API Gateway handles routing, API composition, authentication, rate limiting, and service discovery. It can also perform several supporting tasks such as request/response transformation, caching, and centralized logging.

8. If API Gateway Is a Single Entry Point, How Does It Handle Millions of Requests?

At this point, we know that the API Gateway is the entry point.

This creates an important system-design question:

If millions of requests come through one API Gateway, won't the API Gateway itself become a bottleneck or single point of failure?

The answer is that "single entry point" does not mean there is only one physical API Gateway instance.

The API Gateway is a logical single entry point, not necessarily a single server. In a production system, it is usually deployed across multiple instances, availability zones, and sometimes regions.

  • Multiple API Gateway Instances
    Instead of one gateway handling every request, multiple gateway instances run in parallel. A load balancer distributes incoming requests across these instances, allowing the system to handle much higher traffic. More instances can also be added through auto-scaling when traffic increases.

  • Multiple Availability Zones
    API Gateway instances can be distributed across multiple Availability Zones. If one AZ fails, traffic can be routed to instances in another AZ, improving availability and fault tolerance.

  • Multiple Regions
    For large-scale applications, API Gateway instances can be deployed across multiple geographic regions. Traffic can be distributed across these regions, and if one region becomes unavailable, traffic can be redirected to another region.

  • Load Balancers
    Load balancers distribute incoming requests across multiple API Gateway instances. This prevents a single instance from becoming overloaded and can also stop sending traffic to unhealthy instances.

  • Service Discovery
    After receiving a request, the API Gateway needs to know where the required microservice is running. Service discovery keeps track of available and healthy service instances so that the gateway can route requests to the correct instance.

  • DNS-Based Traffic Distribution
    DNS services such as AWS Route 53 or Azure Traffic Manager can distribute users across different regions or endpoints. This provides an additional layer of traffic distribution before requests reach the API Gateway.

Availability Zone

Now we have solved the problem of distributing traffic across multiple service instances. But another problem appears: what happens if the infrastructure hosting these services fails?

This is where Availability Zones (AZs) come into the picture. A Region contains multiple isolated Availability Zones, and services can be deployed across them. If one Availability Zone fails, traffic can be redirected to healthy instances in another AZ, allowing the application to continue serving users. This provides high availability and fault tolerance within a region.

Multi-AZ API Gateway Architecture

Now we can deploy our API Gateway and microservices across multiple Availability Zones (AZs) within the same region.

Each AZ runs its own API Gateway and service instances. Traffic is distributed across these AZs, so the application does not depend on a single AZ.

If AZ1 fails, traffic can be redirected to AZ2, allowing the application to continue serving requests.

This improves the system's availability and fault tolerance.

Multi-Region Architecture

Multi-AZ protects the application from failures within a region. But if an entire region becomes unavailable, we need another level of redundancy. This is where Multi-Region Architecture is used.

The application can be deployed across multiple regions, with each region having its own API Gateway, Load Balancers, Service Discovery, and microservice instances. Traffic can be distributed between regions using global DNS or traffic-routing services.

If one Availability Zone fails, another AZ within the same region can continue serving traffic. If the entire region fails, traffic can be redirected to another healthy region, allowing the application to continue serving users.

15. DNS-Based Traffic Distribution

Now we have multiple regions.

This creates another question:

If there are multiple API Gateways in different regions, who decides which region receives the client's request?

Now that we have multiple regions, another question arises: how does the system decide which region should handle a client’s request? This is where DNS-based traffic distribution comes into play. Services such as AWS Route 53 and Azure Traffic Manager can direct users to different regions based on configured routing rules.

For example, the traffic manager can consider factors such as latency, geographic location, or compliance requirements when selecting a region. Once a region is selected, the request is sent to the API Gateway in that region. This allows traffic to be distributed across regions and helps the application remain available if a region becomes unavailable.

Complete API Gateway Architecture

Now we can combine everything we have learned.

Remember the responsibilities

ComponentMain Responsibility
DNS / Traffic ManagerChoose the appropriate region
API GatewayRoute request to the correct service
AuthenticationVerify the client/token
Rate LimitingControl request rate
Service DiscoveryFind current service locations
Load BalancerDistribute traffic across service instances
MicroserviceProcess the actual business request

One-Line Mental Model

DNS chooses the region → API Gateway chooses the service → Service Discovery finds its location → Load Balancer chooses the instance → Microservice processes the request.

Interview Answer: What is an API Gateway?

An API Gateway is a single entry point between clients and backend microservices in a microservices architecture.

Instead of the client directly communicating with multiple services, it sends requests to the API Gateway. The gateway identifies which microservice should handle the request and routes it accordingly. For example, a request to /api/order would be routed to the Order Service.

Apart from routing, an API Gateway can also handle common responsibilities such as authentication, rate limiting, API composition, service discovery, caching, and logging.

In a production system, multiple API Gateway instances can run across Availability Zones and regions, so the gateway itself does not become a bottleneck or single point of failure.

Simple Example

Suppose an e-commerce application has:

  • Order Service
  • Payment Service
  • Product Service
Next TopicRate Limiting