Cilium agent — A DaemonSet on every node that loads eBPF programs and enforces policy. It watches the Kubernetes API, manages endpoints and identities on its node and programs the datapath. Docs
Cilium operator — Cluster-wide tasks done once: IPAM allocation, garbage collection. Not in the datapath — if it is briefly down, existing traffic keeps flowing. Docs
IPAM — How Pod IPs are allocated: cluster-pool (default), kubernetes, multi-pool, cloud ENI. In cluster-pool mode the operator hands each node a CIDR from clusterPoolIPv4PodCIDRList. Docs
Routing modes — Encapsulation (VXLAN/Geneve overlay) or native routing. Encapsulation works on any network; native routing needs the underlying network (or BGP/cloud routes) to know Pod CIDRs, but avoids overhead. Docs
Network Policy (18%)
Identity-based security, enforcement modes, rule structure, K8s vs Cilium policies
Security identity — Policy follows labels, not IPs: Pods with the same labels share a numeric identity. The identity travels with the packet, so enforcement does not care how often Pod IPs churn. Docs
CiliumNetworkPolicy — Extends NetworkPolicy with L7 (HTTP, DNS, Kafka) rules, FQDNs and entities. endpointSelector picks the Pods; ingress/egress rules use fromEndpoints, toFQDNs, toEntities (world, host, kube-apiserver) and toPorts.rules.http. Docs
CiliumClusterwideNetworkPolicy — The same policy language, cluster-scoped — including host (node) policies. Useful for baseline rules across all namespaces and for protecting the nodes themselves. Docs
Policy enforcement modes — default, always, never. default: an endpoint becomes default-deny (per direction) once a policy selects it. always: deny unless allowed, even without policies. never: no enforcement. Docs
Service Mesh (16%)
Ingress and Gateway API, mesh use cases, encryption in transit, sidecarless
Gateway API — Cilium implements Gateway and HTTPRoute (and Ingress) with embedded Envoy. Gateway API is role-oriented and more expressive than Ingress: header matching, traffic splitting and cross-namespace routes, standardized. Docs
Sidecarless mesh — L4 in eBPF, L7 in a per-node Envoy — no proxy injected into each Pod. Fewer proxies means less resource overhead and no Pod restarts to join the mesh; sidecar meshes isolate the proxy per workload instead. Docs
Transparent encryption — WireGuard or IPsec encrypts Pod traffic between nodes, no app changes. Enable with encryption.enabled=true and encryption.type=wireguard (or ipsec). `cilium encrypt status` checks it. Docs
Mutual authentication — SPIFFE identities verified between workloads, required per policy rule. A policy rule with authentication.mode: required makes Cilium verify both peers before allowing traffic. Docs
Network Observability (10%)
Hubble, L7 visibility, the Hubble CLI and UI
Hubble — Flow visibility built on Cilium’s eBPF datapath. Every flow with source/destination identities, verdict (FORWARDED/DROPPED) and, with L7 visibility, HTTP and DNS details. Hubble Relay aggregates all nodes. Docs
hubble observe — CLI to filter flows: --namespace, --pod, --verdict DROPPED, --protocol http. The fastest way to see which policy dropped a connection. `cilium hubble port-forward` exposes Relay locally. Docs
Hubble UI — A live service map of who talks to whom. `cilium hubble ui` opens it; flows and verdicts are shown per namespace. Docs
L7 visibility — HTTP methods, paths, status codes and DNS queries in flows. Enabled by L7 rules in a CiliumNetworkPolicy, which route the traffic through the node’s Envoy proxy. Docs
cilium install — Installs Cilium (Helm under the hood) with --set overrides. `cilium install --version 1.x --set kubeProxyReplacement=true`. Helm installs work just as well. Docs
cilium status — Health of the agents, operator, Hubble and managed Pods. `cilium status --wait` blocks until everything is ready. Docs
cilium connectivity test — Deploys test workloads and checks Pod, Service, policy and external connectivity. Run it after install or upgrade to validate the datapath end to end. Docs
cilium config — View or change settings in the cilium-config ConfigMap. `cilium config view`, `cilium config set <key> <value>` (restarts agents to apply). Docs
Cluster Mesh (10%)
Multi-cluster connectivity, service discovery and load balancing
Cluster Mesh — Connects clusters: Pod-to-Pod routing, shared identities and policies. `cilium clustermesh enable` on each cluster, then `cilium clustermesh connect`. Each cluster needs a unique name, ID and non-overlapping Pod CIDRs. Docs
Global service — annotation service.cilium.io/global: "true" load-balances across clusters. Same-named Services in the same namespace on several clusters merge endpoints; service.cilium.io/affinity can prefer local ones. Docs
clustermesh-apiserver — Exposes a cluster’s identities, endpoints and services to its peers. Backed by its own etcd; remote agents connect to it to learn about the other cluster. Docs
eBPF (10%)
eBPF’s role in Cilium, its benefits, eBPF vs iptables
eBPF — Sandboxed programs run inside the Linux kernel, verified before loading. Programs attach to hooks (XDP, tc, sockets, kprobes) and are checked by the kernel verifier so they cannot crash the kernel. Docs
eBPF maps — Kernel key/value stores shared by eBPF programs and user space. Cilium keeps services, endpoints, policies and connection tracking in maps — O(1) lookups instead of long rule chains. Docs
kube-proxy replacement — Service load balancing in eBPF instead of iptables rules. iptables evaluates rules sequentially and rewrites big chains on every change; eBPF hash lookups stay fast as Services grow. Docs
BGP and External Networking (6%)
Egress requirements and connecting clusters to external networks
BGP Control Plane — Peers with routers to advertise Pod CIDRs and LoadBalancer IPs. Configured with CiliumBGPClusterConfig, CiliumBGPPeerConfig and CiliumBGPAdvertisement resources. Docs
Egress gateway — Sends selected Pods’ outbound traffic through specific nodes with a fixed source IP. For firewalls outside the cluster that allow-list IPs: a CiliumEgressGatewayPolicy picks the Pods, destinations and gateway node. Docs
LB IPAM — Assigns IPs to LoadBalancer Services from a CiliumLoadBalancerIPPool. Pair it with BGP or L2 announcements so the outside world can reach those IPs on bare metal. Docs
Practice questions
What role does Cilium play in a Kubernetes cluster?
Answer: A CNI plugin providing networking, security and observability with eBPF. It can also replace kube-proxy and provide service mesh features.
What is a Cilium endpoint?
Answer: A group of containers sharing an IP — usually one per Pod — managed by the agent. Policy is enforced per endpoint.
Where does Cilium store identities and other cluster-wide state by default?
Answer: As Kubernetes custom resources (CRD mode). A kvstore (etcd) can be used instead at very large scale.
Which CRD represents a Cilium security identity?
Answer: CiliumIdentity. Identities are derived from security-relevant labels.
Which CRD represents a node’s Cilium configuration, such as its Pod CIDRs?
Answer: CiliumNode. The operator allocates CIDRs into CiliumNode resources in cluster-pool IPAM.
Which IPAM mode lets Kubernetes itself assign Pod CIDRs to nodes (spec.podCIDR)?
Answer: kubernetes. cluster-pool (the default) has the Cilium operator allocate them instead.
Which IPAM mode allocates Pod IPs directly from AWS VPC networking?
Answer: eni. Pods get routable VPC IPs from ENIs attached to the node.
In tunnel (encapsulation) mode, which protocols can Cilium use?
Answer: VXLAN; Geneve. VXLAN is the default tunnel protocol.
What is required for native routing mode?
Answer: The network must route Pod CIDRs between nodes (e.g. via BGP or cloud routes). Set routingMode=native and ipv4NativeRoutingCIDR.
What does the Cilium agent use to receive Kubernetes state?
Answer: It watches the Kubernetes API (Pods, Services, policies, nodes). It translates that state into eBPF programs and maps.
What does the cilium-envoy component do?
Answer: Runs the Envoy proxy used for L7 policy, Ingress/Gateway API and visibility. It runs per node (as a DaemonSet or embedded in the agent).
What is Cilium’s health checking used for?
Answer: Probing connectivity between nodes and endpoints (cilium-health). `cilium-health status` reports node-to-node reachability.
Which component handles cluster-wide garbage collection of stale identities?
Answer: The Cilium operator. The operator handles cluster-wide tasks once rather than on every node.
What does Cilium’s kube-proxy replacement need from the kernel?
Answer: eBPF features such as socket-level and tc/XDP hooks in a recent kernel. Check the system requirements page for minimum kernel versions.
What is "socket-level load balancing" in Cilium?
Answer: Translating a Service IP to a backend at connect() time, before packets are created. It avoids per-packet NAT for in-cluster traffic.
Which datapath mode can Cilium use to bypass the host network stack for Pod traffic on supported setups?
Answer: netkit or veth with eBPF host-routing. eBPF host-routing avoids much of the upper network stack for performance.
Which CiliumNetworkPolicy field selects the Pods a policy applies to?
Which field allows ingress from Pods with label app=api?
Answer: ingress[].fromEndpoints with matchLabels app: api. fromEndpoints selects sources by label.
Which policy entity represents everything outside the cluster?
Answer: world. Other entities include host, remote-node, kube-apiserver, cluster and all.
How do you allow egress to api.github.com by name?
Answer: egress toFQDNs with matchName: api.github.com, plus a DNS rule allowing the lookup. Cilium learns the IPs from DNS responses it observes.
Why must a toFQDNs policy also allow DNS to kube-dns with an L7 DNS rule?
Answer: Cilium needs to see the DNS answers to map names to IPs. The DNS proxy records name-to-IP mappings for policy.
Which L7 rule allows only GET /public on port 80?
Answer: toPorts with ports 80/TCP and rules.http [{method: GET, path: /public}]. Requests that fail L7 rules are answered with HTTP 403 by Envoy.
What happens to a request denied by an L7 HTTP rule?
Answer: The proxy returns HTTP 403 Access Denied. L3/L4 denials drop packets; L7 denials return a response.
How do you express a deny rule in Cilium policy?
Answer: ingressDeny / egressDeny sections. Deny rules take precedence over allow rules.
Which policy type can protect the nodes themselves (host firewall)?
Answer: A CiliumClusterwideNetworkPolicy with a nodeSelector. Host policies need the host firewall feature enabled.
In policyEnforcementMode "always", what happens to an endpoint with no policy selecting it?
Answer: All its traffic is denied. In "default" mode it would be allowed until a policy selects it.
Do Kubernetes NetworkPolicies work with Cilium?
Answer: Yes — Cilium enforces standard NetworkPolicy alongside its own CRDs. Both are translated into the same eBPF policy maps.
Which label is excluded from identity by default because it changes often?
Answer: pod-template-hash (and similar non-security labels). Identity-relevant labels can be configured to avoid identity churn.
Which rule allows traffic from Pods in namespace monitoring?
Answer: fromEndpoints matchLabels k8s:io.kubernetes.pod.namespace: monitoring. Namespace is exposed as a special label in Cilium policy.
Which Cilium Helm value enables Gateway API support?
Answer: gatewayAPI.enabled=true. The Gateway API CRDs must be installed beforehand; kubeProxyReplacement is required too.
Which Cilium Helm value enables its Ingress controller?
Answer: ingressController.enabled=true. Ingress resources then use ingressClassName: cilium.
What is a benefit of Gateway API over Ingress?
Answer: Portable, standard features like header matching and traffic splitting without vendor annotations. It also separates roles: infrastructure vs application owners.
How does Cilium implement traffic splitting for an HTTPRoute?
Answer: Envoy applies the backendRefs weights. Weights are honoured per request.
Which encryption option is generally simpler to operate in Cilium?
Answer: WireGuard (keys are managed automatically per node). IPsec requires creating and rotating a key Secret.
Which command shows whether transparent encryption is active?
Answer: cilium encrypt status. It reports the mode (WireGuard/IPsec) and peers.
What is a sidecarless service mesh advantage?
Answer: No proxy injected per Pod, so lower resource use and no Pod restarts to join. Sidecars isolate per workload; per-node proxies trade that for efficiency.
Which use case is typical for a service mesh?
Answer: mTLS between services; Traffic shifting for canaries; Retries and timeouts. Mesh features operate on service-to-service traffic.
What does Cilium mutual authentication use for identities?
Answer: SPIFFE identities issued via SPIRE. Policies with authentication.mode: required check them.
Which Kubernetes resource does Cilium use to represent Gateway API listeners?
Answer: Gateway. Cilium implements the upstream Gateway API resources.
Which GatewayClass controllerName does Cilium use?
Answer: io.cilium/gateway-controller. Gateways reference a GatewayClass handled by that controller.
Encryption in transit protects against…
Answer: Eavesdropping on traffic between nodes. It protects data on the wire, not workloads themselves.
Which component aggregates flows from every node’s Hubble server?
Answer: Hubble Relay. The CLI and UI usually talk to Relay.
Which command shows only dropped flows for namespace shop?
Answer: hubble observe --namespace shop --verdict DROPPED. Filters can be combined: --pod, --protocol, --port, --to-fqdn.
How do you enable Hubble with the Cilium CLI?
Answer: cilium hubble enable (add --ui for the UI). Helm users set hubble.enabled, hubble.relay.enabled and hubble.ui.enabled.
Which flows show HTTP status codes in Hubble?
Answer: Flows for traffic covered by L7 visibility (policy or visibility annotations). L3/L4 flows carry no HTTP details.
What can Hubble export to Prometheus?
Answer: Flow-based metrics such as dns, drop, tcp, flow and http. Enable them with hubble.metrics.enabled.
Which command streams raw datapath events on a single node for debugging?
Answer: cilium-dbg monitor (inside the agent Pod). It shows drops, traces and policy verdicts from that node.
Which flag makes hubble observe follow new flows continuously?
Answer: --follow (-f). Without it, the CLI prints recent flows and exits.
A flow shows verdict DROPPED with reason "Policy denied". What next?
Answer: Check which policies select the source and destination endpoints. cilium-dbg policy get and Hubble’s policy fields help find the rule.
Which command installs Cilium with the Cilium CLI?
Answer: cilium install. It auto-detects many platform settings; --set passes Helm values.
Which command waits until Cilium is ready?
Answer: cilium status --wait. It reports agent, operator and Hubble status.
Which command upgrades Cilium with the CLI?
Answer: cilium upgrade --version <v>. Follow the upgrade guide: upgrade one minor version at a time with pre-flight checks.
Where does the Cilium agent read most of its settings from?
Answer: The cilium-config ConfigMap in kube-system. Helm values render into this ConfigMap.
Which command runs the end-to-end connectivity checks?
Answer: cilium connectivity test. It deploys test workloads into a cilium-test namespace.
After cilium config set, what is needed for agents to pick up the change?
Answer: Restarting the agents (the CLI does this by default). Many settings are only read at agent start.
How can you install Cilium besides the Cilium CLI?
Answer: With the cilium Helm chart. The CLI itself uses the Helm chart under the hood.
Which setting must be true to remove kube-proxy and let Cilium handle Services?
Answer: kubeProxyReplacement=true (with k8sServiceHost/Port set). Cilium must reach the API server directly when kube-proxy is gone.
Which settings must be unique per cluster in a Cluster Mesh?
Answer: cluster.name; cluster.id. Pod CIDRs must also not overlap.
Which command connects two clusters after enabling Cluster Mesh in both?
Answer: cilium clustermesh connect --context c1 --destination-context c2. cilium clustermesh status --wait checks the result.
How are the clustermesh-apiservers exposed to the other clusters?
Answer: Through a Service such as LoadBalancer or NodePort. Remote agents need to reach it over the network.
A Service is annotated global and has endpoints in two clusters. Which annotation prefers local endpoints?
Answer: service.cilium.io/affinity: local. Remote endpoints are then used only when no local ones are healthy.
What does service.cilium.io/shared: "false" do on a global Service?
Answer: This cluster uses remote endpoints but does not share its own. Useful to consume a global service without exporting local backends.
Can network policies select endpoints in another cluster of the mesh?
Answer: Yes, using the io.cilium.k8s.policy.cluster label. Identities are shared across the mesh.
What is a typical Cluster Mesh use case?
Answer: High availability: fail over a service to another cluster. Also shared services and splitting stateful/stateless across clusters.
Which command shows Cluster Mesh connectivity status?
Answer: cilium clustermesh status. It shows connected clusters and global service counts.
What does the eBPF verifier guarantee before a program loads?
Answer: It terminates and only accesses memory safely. Unsafe programs are rejected by the kernel.
Which eBPF hook runs earliest on packet receive, in the NIC driver?
Answer: XDP. XDP is used for very fast load balancing and filtering.
Why can eBPF programs be updated without restarting applications?
Answer: They are loaded into the running kernel and attached to hooks dynamically. No kernel modules or restarts are needed.
How does Cilium track connections for NAT and policy?
Answer: In its own eBPF connection-tracking maps. Map sizes are tunable for busy nodes.
Which is a key benefit of eBPF-based networking over iptables at scale?
Answer: Lookup cost does not grow linearly with the number of Services. iptables evaluates rule chains sequentially.
What is an eBPF map?
Answer: A kernel data structure shared between eBPF programs and user space. Hash, array, LRU and other map types exist.
Which non-networking use of eBPF does Tetragon provide?
Answer: Security observability and runtime enforcement from kernel events. Tetragon is a Cilium sub-project.
Do eBPF programs require loading kernel modules?
Answer: No — they run in the kernel’s built-in eBPF virtual machine. This is part of what makes eBPF safe to deploy.
Which resource defines BGP peers and the local ASN in Cilium’s BGP control plane (v2 API)?
Answer: CiliumBGPClusterConfig (with CiliumBGPPeerConfig). Advertisements are configured with CiliumBGPAdvertisement.
What can Cilium advertise over BGP?
Answer: Pod CIDRs; LoadBalancer Service IPs. Advertising Service IPs makes LoadBalancer Services reachable on bare metal.
Without BGP, how can Cilium announce LoadBalancer IPs on a flat L2 network?
Answer: L2 announcements (ARP/NDP). A node answers ARP requests for the Service IP.
Which resource gives Pods a stable egress source IP?
Answer: CiliumEgressGatewayPolicy. Traffic is SNATed on the chosen gateway node.
Why might egress traffic need masquerading (SNAT)?
Answer: Pod IPs are not routable outside the cluster. Native routing setups may disable masquerading when Pod IPs are routable.
Which Cilium component runs on every node and enforces policy in the datapath?
Answer: The Cilium agent (DaemonSet). The agent loads eBPF programs per node; the operator handles cluster-wide chores and is not in the datapath.
Your underlying network cannot route Pod CIDRs. Which routing mode works out of the box?
Answer: Encapsulation (VXLAN or Geneve). Overlay encapsulation only needs node-to-node connectivity; native routing needs the network to know Pod CIDRs.
How does Cilium identify the source of traffic for policy decisions?
Answer: By a security identity derived from the Pod’s labels. Identity-based enforcement survives Pod IP churn and scales better than IP lists.
Which can a CiliumNetworkPolicy do that a standard Kubernetes NetworkPolicy cannot?
Answer: Allow only GET /api/* (L7 HTTP rules); Allow egress to a DNS name with toFQDNs. Both select Pods and ports; L7 rules and FQDN-based egress are Cilium extensions.
In the default enforcement mode, when does a Pod become default-deny for ingress?
Answer: As soon as any policy with an ingress section selects it. default mode isolates an endpoint (per direction) once a policy selects it; "always" would deny from the start.
Which Cilium option encrypts Pod-to-Pod traffic between nodes without app changes?
Answer: Transparent encryption with WireGuard or IPsec. Encryption is applied by the datapath, so applications need no TLS of their own for it.
What is a key benefit of Gateway API over Ingress?
Answer: Standardized, role-oriented resources with richer routing (header matches, traffic splitting). Ingress relies on vendor annotations for anything advanced; Gateway API makes it portable.
A connection is refused and you suspect a policy drop. Fastest check?
Answer: hubble observe --verdict DROPPED --pod <pod>. Hubble shows each flow with its verdict and the identities involved.
After installing Cilium, how do you validate Pod, Service and policy connectivity end to end?
Answer: cilium connectivity test. The connectivity test deploys probe workloads and runs a suite of checks.
What makes a Service load-balance across all clusters in a Cluster Mesh?
Answer: The annotation service.cilium.io/global: "true" on same-named Services in each cluster. Global services merge endpoints from every cluster that has that Service.
Why does eBPF-based service load balancing scale better than iptables?
Answer: Hash-table map lookups instead of evaluating long sequential rule chains. iptables cost grows with the number of rules; eBPF maps give near-constant lookups.
An external firewall only allows a fixed source IP. How do you route certain Pods’ egress through it?
Answer: A CiliumEgressGatewayPolicy sending that traffic via a gateway node. The egress gateway SNATs the selected traffic to a predictable IP on the chosen node.