Istio Service Mesh in Practice: What It Gives You and Where It Bites
A service mesh moves the hard parts of service-to-service networking - encryption, retries, timeouts, traffic shifting and telemetry - out of your application code and into the platform. Istio is the most widely used mesh on Kubernetes, and when it’s adopted deliberately it removes a lot of bespoke plumbing. Adopted all at once, it can also add a layer of complexity your team isn’t ready to debug.
This post covers what Istio gives you, the features worth turning on first, and the gotchas we’ve hit running it in production.
How Istio works
Istio has a control plane (istiod), which turns your Istio resources
into proxy configuration, and a data plane of Envoy proxies that
actually carry the traffic. There are two data-plane modes:
- Sidecar mode - an Envoy proxy is injected into every pod. It’s the long-standing model with the richest feature set, but every pod pays the CPU and memory cost of its own proxy.
- Ambient mode - a per-node
ztunnelhandles mTLS and layer-4 traffic, and optional waypoint proxies add layer-7 features only where you need them. No sidecars, far lower overhead, and no pod restarts to join the mesh. Ambient reached general availability in Istio 1.24.
For new meshes, ambient is worth evaluating first. Sidecar mode is still the safer choice if you depend on features ambient doesn’t support yet.
Start with mutual TLS
Encrypting and authenticating every service-to-service call is often the single biggest win - and a common compliance requirement. Istio issues workload certificates and rotates them automatically.
Roll it out in two steps. First run in the default permissive mode, where services accept both plaintext and mTLS, while you confirm every workload has joined the mesh. Then enforce strict mode, namespace by namespace:
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
name: default
namespace: payments
spec:
mtls:
mode: STRICT
With identities in place you can add AuthorizationPolicy resources to control which services may call each other - network policy based on workload identity rather than IP addresses.
Traffic management
Istio’s traffic features make releases safer without touching application
code. A DestinationRule defines versions (subsets) of a service and
connection behaviour; a VirtualService decides how requests are routed.
A canary release sending 10% of traffic to a new version, with retries and a timeout:
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: checkout
spec:
hosts:
- checkout
http:
- route:
- destination:
host: checkout
subset: v1
weight: 90
- destination:
host: checkout
subset: v2
weight: 10
retries:
attempts: 3
perTryTimeout: 2s
timeout: 10s
And the matching DestinationRule, with outlier detection - Istio’s
circuit breaker - ejecting instances that keep returning errors:
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: checkout
spec:
host: checkout
trafficPolicy:
outlierDetection:
consecutive5xxErrors: 5
interval: 30s
baseEjectionTime: 60s
subsets:
- name: v1
labels:
version: v1
- name: v2
labels:
version: v2
Be careful with retries: a retry policy on every hop of a deep call chain can multiply load on a struggling service. Keep retry counts low and make sure the operations being retried are safe to repeat.
Where it bites
Sidecar overhead adds up. Every sidecar holds configuration for every
service in the mesh by default. In large clusters that means significant
memory per pod. A Sidecar resource scopes each namespace’s proxies to just
the services they talk to:
apiVersion: networking.istio.io/v1
kind: Sidecar
metadata:
name: default
namespace: payments
spec:
egress:
- hosts:
- "./*"
- "istio-system/*"
Startup ordering. If your application starts before its sidecar is ready,
its first outbound calls fail. Setting holdApplicationUntilProxyStarts: true
in the proxy config makes the app container wait for the proxy.
Upgrades. Upgrading the control plane in place is risky. Use
revision-based (canary) upgrades: install the new istiod revision
alongside the old one, move namespaces across one at a time, and roll back
by switching the label if something goes wrong.
Egress gateways and TLS. Routing outbound traffic through an egress
gateway gives you a single, auditable exit point - but TLS routing has sharp
corners. We hit one where combining hosts: ["*"] with
tls.mode: ISTIO_MUTUAL and SNI-based routing (sniHosts) silently failed to
install the gateway’s :443 listener. ISTIO_MUTUAL wants to terminate TLS
before routing, while sniHosts needs to read the SNI from the ClientHello
before termination - and Envoy can’t build a filter chain that does both.
The fix is to replace the wildcard with an explicit list of hostnames, or
switch to PASSTHROUGH and handle mTLS at the sidecar. We wrote it up in a
discussion on the Istio repo.
The lesson generalises: when Istio config “applies cleanly” but traffic
doesn’t flow, check what Envoy actually received. istioctl analyze catches
many misconfigurations up front, and istioctl proxy-config listeners /
routes show the config each proxy is really running.
Observability comes for free - mostly
Because every request passes through Envoy, the mesh emits consistent request metrics (rate, errors, latency) for every service, whatever language it’s written in. Feed them into Prometheus or a platform like Datadog and you get a service map and golden signals with no code changes. Distributed tracing still needs your applications to propagate trace headers - the mesh can’t stitch spans together on its own.
For more on turning that telemetry into useful alerts, see our post on Datadog on Kubernetes.
Is a service mesh right for you?
A mesh earns its keep when you have many services, a requirement for encryption in transit, or releases that need fine-grained traffic control. If you run a handful of services, Kubernetes network policies and a good ingress controller may be all you need.
New to Istio? The Istio Masterclass talk is a good introduction. If you’re planning a mesh rollout or fighting one in production, get in touch - it’s one of the services we offer.