ScreenshotNeo

BlogEngineering

Understanding Kubernetes Architecture

Learn how Kubernetes control planes, nodes, controllers, scheduling, networking and security work together to run applications reliably.

By the ScreenshotNeo team1 October 20269 min read

Short answer: A Kubernetes cluster has a control plane and one or more worker nodes. The control plane exposes the API, stores desired and observed state, schedules Pods and runs reconciliation controllers. Worker nodes run Pods through the kubelet and a container runtime, while kube-proxy or an equivalent network plugin provides Service networking. Kubernetes continuously compares actual state with declared state and makes changes until they match.

This guide explains each component, the request path from an object submission to a running container, deployment and availability choices, communication boundaries, ports, failure behavior and operational troubleshooting.

1. Architecture at a glance

The control plane makes cluster-wide decisions. Nodes provide compute capacity. You normally interact with the API server rather than contacting schedulers, controllers or kubelets directly.

Layer Component Primary responsibility
Control plane kube-apiserver Exposes the Kubernetes HTTP API and acts as the front door for users, nodes and controllers.
Control plane etcd Stores serialized Kubernetes API objects in a consistent, highly available key-value store.
Control plane kube-scheduler Chooses a suitable node for each unscheduled Pod.
Control plane kube-controller-manager Runs built-in control loops that reconcile resources.
Control plane cloud-controller-manager Runs cloud-specific controllers when a cloud provider integration is used.
Node kubelet Ensures containers described by PodSpecs are running and healthy.
Node Container runtime Downloads images and runs containers for Pods.
Node kube-proxy (optional) Maintains node networking rules for Services; a network plugin can provide equivalent behavior.
Add-ons DNS, monitoring, logging, dashboard Extend the core cluster with discovery and operational capabilities.

See the Kubernetes cluster architecture documentation for the component model.

2. The control plane

kube-apiserver: the API hub

The API server is the component that exposes the Kubernetes API. kubectl, controllers, kubelets and external automation submit HTTP requests to it. Authentication, authorization and admission processing occur at this boundary before an object is accepted.

The API server is deliberately the hub of the normal communication pattern. Nodes and Pods use the API server for control-plane communication, while other control-plane components are not intended to expose remote services directly.

etcd: the durability boundary

etcd stores Kubernetes API data, including desired configuration, status and metadata. A successful API write is not the same as a workload being ready: controllers, the scheduler and kubelets still have to act on the stored object.

Protect etcd as control-plane state. Restrict network access, use encryption and authentication as appropriate, monitor disk latency and maintain tested backups. A cluster can lose scheduling temporarily and recover; loss or corruption of etcd data threatens the cluster’s source of truth.

kube-scheduler: placement decisions

The scheduler watches for Pods without an assigned node. It filters and scores candidates using resource requests, hardware and software constraints, affinity and anti-affinity, data locality, inter-workload interference and deadlines. It records the selected node through the API server; it does not start the container.

Controllers: continuous reconciliation

Controllers are control loops that watch cluster state and make or request changes when observed state differs from desired state. A Deployment controller creates or replaces ReplicaSets, a ReplicaSet controller creates Pods, and a Node controller reacts to node health signals. Controllers communicate through the API server, which lets many focused controllers cooperate and also permits custom controllers outside the built-in control plane.

cloud-controller-manager

When a cloud integration is enabled, this optional component handles provider-specific logic such as node lifecycle, routes and load balancers. Its exact responsibilities depend on the provider and integration.

3. Worker nodes and Pods

kubelet

The kubelet is the primary node agent. It receives PodSpecs through the control plane, asks the container runtime to create the declared containers and reports status and health. It ignores containers it did not create, so it is not a general-purpose process supervisor for arbitrary host workloads.

Container runtime

The runtime pulls images, creates containers, configures namespaces and cgroups and reports container state to the kubelet. Kubernetes uses a runtime that implements the supported Container Runtime Interface.

kube-proxy and the network plugin

Services provide a stable virtual destination for changing Pod endpoints. kube-proxy usually maintains node-level forwarding rules for that abstraction. Some network plugins implement equivalent Service proxying, making kube-proxy unnecessary. The plugin also supplies Pod-to-Pod networking and commonly enforces network policy.

Node health and eligibility

Nodes report status and heartbeats, including Lease objects in the kube-node-lease namespace. The control plane only considers a node eligible when its object and required services are healthy. If heartbeats stop, controllers can mark the node unhealthy and workloads may be rescheduled according to their controllers and disruption policies.

4. From declarative object to running Pod

  1. A user or automation client submits a Deployment, Job, Pod or another object with kubectl or an API client.
  2. The API server authenticates and authorizes the request, runs admission, then persists the accepted object in etcd.
  3. Controllers watch the API and create related objects until the desired relationships exist.
  4. The scheduler notices each Pod without a node and selects a feasible node.
  5. The scheduler writes the binding through the API server.
  6. The selected node’s kubelet observes the PodSpec and asks the container runtime to pull images and start containers.
  7. The kubelet performs probes and reports status. Controllers replace failed or missing replicas when the declared workload requires it.
  8. Service networking routes traffic to ready Pod endpoints through kube-proxy or the network plugin.

This is a control loop, not a one-time transaction. If a container exits, a node fails or an operator changes the declaration, components continue reconciling toward the new desired state.

Minimal runnable example

kubectl apply -f - <<'YAML'
apiVersion: apps/v1
kind: Deployment
metadata:
  name: web
spec:
  replicas: 3
  selector:
    matchLabels:
      app: web
  template:
    metadata:
      labels:
        app: web
    spec:
      containers:
      - name: web
        image: nginx:1.27
        ports:
        - containerPort: 80
        resources:
          requests:
            cpu: 100m
            memory: 128Mi
          limits:
            cpu: 500m
            memory: 256Mi
---
apiVersion: v1
kind: Service
metadata:
  name: web
spec:
  selector:
    app: web
  ports:
  - port: 80
    targetPort: 80
YAML

kubectl get pods -o wide
kubectl describe deployment web
kubectl get endpointslices -l kubernetes.io/service-name=web

5. Communication paths and security boundaries

The documented hub-and-spoke pattern terminates node and Pod API usage at the API server. Secure the API endpoint with TLS, authenticate every client and apply least-privilege authorization.

The API server also reaches kubelet endpoints for logs, attach and port-forward. On untrusted networks, configure deliberate kubelet authentication and authorization and verify kubelet certificates from the API server. API-server proxy paths to nodes, Pods and Services have different default protection characteristics; review each path before exposing it outside trusted networks.

Connection Purpose Boundary to review
Operator or CI → API server Submit and read API objects TLS, authentication, RBAC and admission.
Node → API server Watch PodSpecs and publish status Node credentials and API authorization.
API server → kubelet Logs, exec, attach and port-forward Kubelet TLS verification and kubelet authn/authz.
Pod → Pod or Service Application traffic Network plugin, NetworkPolicy and Service exposure.
Control plane → etcd Read and write cluster state Client certificates, firewalling and backup access.

Default ports

Defaults can be overridden; firewall rules must follow the actual configuration.

Component Default port
API server TCP 6443
etcd client and peer traffic TCP 2379–2380
kubelet TCP 10250
scheduler TCP 10259
controller-manager TCP 10257
kube-proxy metrics or health endpoint TCP 10256
NodePort range TCP/UDP 30000–32767

Reference: Kubernetes ports and protocols.

6. Deployment and availability choices

Self-managed control plane

Components can run as services on dedicated machines, as static Pods managed by a kubelet or in a self-hosted arrangement. You own patching, certificates, etcd backups, monitoring, capacity planning and incident response.

Managed Kubernetes

A managed Kubernetes service abstracts some control-plane operations. Confirm what the provider operates, which versions and regions are available, how upgrades work, how control-plane access is secured and what backup and support responsibilities remain yours.

Single versus distributed control planes

A single control-plane machine is simpler and cheaper but has a larger failure domain. Production clusters commonly distribute control-plane components and run multiple worker nodes for fault tolerance. More machines add coordination, network and operational complexity; they do not remove the need for tested backups and failure procedures.

Decision Question
Ownership Who patches, upgrades, monitors and backs up the control plane?
Failure tolerance How many control-plane or worker failures must the cluster survive?
Reachability Which networks can reach the API server and kubelets?
Customization Do you need custom schedulers, API extensions or controllers?
Cost and staffing Can your team operate the control plane, or is provider management worth the fee?

7. Performance, reliability and cost considerations

  • API and etcd: API request volume, object size and etcd disk latency affect control-plane responsiveness. Avoid writing status or events in tight loops.
  • Scheduling: Accurate resource requests improve placement. Excessive constraints, affinity rules or thousands of pending Pods increase scheduling work.
  • Node capacity: Leave headroom for system Pods, image pulls, upgrades and failure recovery. Limits control bursts; requests drive placement.
  • Networking: Service and overlay behavior depends on the chosen network plugin. Measure your own traffic patterns instead of assuming a universal throughput figure.
  • Reliability: Use multiple replicas, readiness probes, disruption budgets and spread rules where appropriate. Back up etcd and rehearse restore procedures.
  • Cost: Kubernetes has no universal architecture cost or performance number. Cost depends on control-plane ownership, worker count, instance sizes, storage, network traffic, observability and staffing.

8. Troubleshooting checklist

Symptom Likely cause Checks and fix
Pod remains Pending No feasible node, insufficient resources or an unsatisfied constraint. Run kubectl describe pod NAME; inspect scheduler events, requests, taints, tolerations, affinity and available capacity.
Pod is scheduled but not Running Image pull, runtime, volume or security failure. Inspect kubectl describe pod and container events; verify image name, registry credentials, storage and node runtime logs.
Deployment has fewer ready replicas Readiness probe failure, crash loop or unavailable node. Check kubectl get pods, probe configuration and container logs. Fix the application or probe before increasing replicas.
Service has no traffic Selector mismatch, no ready endpoints or network-policy block. Compare Service selectors with Pod labels; inspect EndpointSlices and NetworkPolicy objects.
Node shows NotReady Kubelet stopped, heartbeat path failed, resource pressure or runtime error. Inspect Node conditions and events, then kubelet and runtime logs on the node. Restore connectivity or capacity.
kubectl cannot connect Wrong kubeconfig, expired credentials, unreachable API endpoint or firewall rule. Run kubectl config current-context, verify the server address and credentials, then test API reachability and TLS.
Logs or exec fail API-server-to-kubelet TLS or authorization misconfiguration. Verify kubelet certificates, API-server kubelet CA settings and kubelet authorization rules.
Control plane is slow etcd disk latency, overloaded API server or excessive watches and writes. Review API and etcd metrics, reduce noisy clients, protect disk I/O and scale the control plane according to measured demand.

9. Practical inspection commands

kubectl cluster-info
kubectl get --raw='/readyz?verbose'
kubectl get nodes -o wide
kubectl describe node NODE_NAME
kubectl get events --sort-by=.lastTimestamp
kubectl get pods -A -o wide
kubectl get componentstatuses

componentstatuses is retained for compatibility on some versions and may not provide a complete health picture. Prefer the version-specific health endpoints, component metrics and events documented for your distribution.

10. Or skip the browser setup

When you publish architecture documentation, you may need clean screenshots of rendered diagrams, runbooks or dashboards. ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP or PDF, with options for full-page capture, element selectors, custom CSS and JavaScript, waiting rules, device presets, headers, cookies, geolocation, caching, signed links, asynchronous jobs and bulk capture.

Use the do-it-yourself browser method when you need full control over a local rendering environment. For hosted pages, this one-call option avoids browser setup:

ScreenshotNeo API documentation

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server lets AI agents use take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

11. FAQ

Is Kubernetes a virtual machine manager?

No. Kubernetes schedules and manages container workloads on physical or virtual machines; the underlying infrastructure remains the responsibility of the node or cloud platform.

Can a cluster run without kube-proxy?

Yes. A network plugin can provide equivalent Service proxying. Confirm the behavior and support model of the plugin you choose.

Does the scheduler start containers?

No. It selects a node and records the assignment. The kubelet and container runtime start and monitor the containers.

Why is etcd so important?

It stores the API-object state that represents the cluster’s desired and observed configuration. Protect it with access controls, backups and tested recovery.

Should control-plane nodes run application Pods?

Development clusters may share capacity. Production environments often dedicate nodes to control-plane duties to reduce interference and preserve recovery capacity.

Where should I start when learning a new cluster?

Trace one Deployment: inspect its API object, ReplicaSet, Pods, scheduler events, node assignment, kubelet status and Service EndpointSlices. This follows the real control loop end to end.