Kubernetes Networking Deep Dive: 16 Diagrams Explain Underlay and Overlay Models
This article provides a comprehensive analysis of Kubernetes networking models, covering underlay (Flannel host-gw, Calico BGP) and overlay (Flannel VxLAN, Calico IPIP/VxLAN, Weave fastdp) approaches, with 16 diagrams illustrating CNI plugins, tunneling protocols, and hardware acceleration technologies like IPVLAN, MACVLAN, and SR-IOV.
Overview
This article explores the Kubernetes network model and analyzes various network models.
Underlay Network Model
What is Underlay Network
The Underlay Network refers to the network device infrastructure such as switches, routers, and DWDM linked by network media to form a physical network topology responsible for packet transmission between networks.
The underlay network can be Layer 2 or Layer 3; a typical Layer 2 example is Ethernet, and a typical Layer 3 example is the Internet.
Technology operating at Layer 2 is vlan, while Layer 3 technologies consist of protocols such as OSPF and BGP.
Underlay Network in Kubernetes
In Kubernetes, a typical example of underlay network is using the host as a router device, where Pod networks learn routing entries to achieve cross-node communication.
Typical implementations under this model include
flannel host-gwmode and
calico BGPmode.
flannel host-gw
In flannel host-gw mode, each Node must be in the same Layer 2 network, and the Node acts as a router. Cross-node communication uses routing tables, simulating the network as an underlay network.
Notes: Because routing is used, the cluster CIDR must be at least /16 to ensure cross-node Nodes act as one network and same-node Pods as another. Otherwise, routing tables in the same network cause unreachable networks.
Calico BGP
BGP(Border Gateway Protocol) is a decentralized autonomous routing protocol. It maintains IP routing tables or prefix tables to achieve reachability between AS (Autonomous Systems) and belongs to the vector routing protocol category.
Unlike flannel, Calico provides a BGP network solution. The network model is similar to Flannel host-gw, but the software architecture differs: flannel uses the flanneld process to maintain routing information, while Calico includes multiple daemons. The Bird process acts as a BGP client and route reflector ( Router Reflector). The BGP client obtains routes from Felix and distributes them to other BGP Peer s, while the reflector optimizes the number of BGP connections within an AS. Typically, the RR is a real routing device, and Bird works as a BGP client.
IPVLAN & MACVLAN
IPVLANand MACVLAN are NIC virtualization technologies. The difference: IPVLAN allows a physical NIC to have multiple IP addresses with all virtual interfaces sharing the same MAC address; MACVLAN allows a single NIC to have multiple MAC addresses, and virtual NICs may have no IP address.
Because they are NIC virtualization rather than network virtualization, they essentially belong to Overlay network. Their key advantage in virtualized environments is flattening Pod networks to the same level as Node networks, providing higher performance and lower latency. The network model corresponds to the second mode in the diagram below.
Virtual Bridge: Creates a veth pair, one end in the container, the other in the host root namespace. Container packets enter the host network stack via the bridge, and packets destined for the container enter via the bridge.
Multiplexing: Uses an intermediate network device exposing multiple virtual NIC interfaces. Container NICs connect to this device, and packets are directed by MAC/IP addresses.
Hardware Switching (SR-IOV): Assigns a virtual NIC to each Pod, making Pod-to-Pod connections nearly equivalent to physical machine communication. Most modern NICs support SR-IOV, which virtualizes a single physical NIC into multiple VF interfaces, each with a separate virtual PCIe channel sharing the physical NIC's PCIe channel.
In Kubernetes, typical CNIs using the IPVLAN model include multus and danm.
multus
multusis an Intel open-source CNI solution that combines traditional cni with multus and provides an SR-IOV CNI plugin enabling K8s Pods to connect to SR-IOV VFs, leveraging IPVLAN/MACVLAN functionality.
When a new Pod is created, the SR-IOV plugin configures the VF, moves it to the new CNI namespace, sets the interface name per the CNI config "name" option, and brings the VF up.
The diagram below shows a Multus and SR-IOV CNI network environment with a three-interface Pod. eth0 is the flannel network plugin, serving as the Pod's default network.
VF is an instantiation of the host physical port ens2f0 (Intel X710-DA4). The Pod-side VF interface is named south0.
Another VF instantiates from host port ens2f1 (another Intel X710-DA4 port). The Pod-side interface is north0, bound to the DPDK driver vfio-pci.
Notes: Terminology NIC : network interface card SR-IOV : single root I/O virtualization, hardware function allowing VMs to share PCIe devices VF : Virtual Function, based on PF, shares physical resources with PF or other VFs PF : PCIe Physical Function, has full control over PCIe resources DPDK : Data Plane Development Kit
Alternatively, the host interface can be moved directly into the Pod's network namespace (the interface must exist and cannot be the same as the default network interface). This places the Pod network on the same plane as the Node network in a standard NIC environment.
danm
DANMis a Nokia open-source CNI project aimed at bringing carrier-grade networking to Kubernetes. Like multus, it provides SR-IOV/DPDK hardware technology and supports IPVLAN.
Overlay Network Model
What is Overlay
An overlay network uses network virtualization technology to build a virtual logical network on top of the underlay network without modifying the physical network architecture. Essentially, overlay network uses one or more tunneling protocols ( tunneling) to encapsulate packets for transport from one network to another; tunneling protocols focus on packets (frames).
Common Network Tunneling Technologies
Generic Routing Encapsulation (GRE): Encapsulates IPv4/IPv6 packets into another protocol's packets, typically operating at Layer 3.
VxLAN (Virtual Extensible LAN): A simple tunneling protocol that encapsulates Layer 2 Ethernet frames into Layer 4 UDP packets, using 4789 as the default port. VxLAN extends VLAN from 4096 (12-bit VLAN ID) to 16 million (24-bit VN·ID) logical networks.
Typical implementations in the overlay model include flannel and
calico VxLANand IPIP modes.
IPIP
IP in IPis also a tunneling protocol, similar to VxLAN, implemented via Linux kernel encapsulation. IPIP requires the kernel module ipip.ko. Check if loaded with lsmod | grep ipip; load with modprobe ipip.
In Kubernetes, IPIP and VxLAN are similar, both using network tunneling. The difference: VxLAN is essentially a UDP packet, while IPIP encapsulates the packet within its own packet.
Notes: Public clouds may disallow IPIP traffic, e.g., Azure.
VxLAN
In Kubernetes, both flannel and calico implement VxLAN using Linux kernel encapsulation. Linux support for VxLAN is relatively recent: Stephen Hemminger merged the work in 2012, appearing in kernel 3.7.0. For stability and features, some software recommends kernel 3.9.0 or 3.10.0+ for VxLAN.
In a Kubernetes VxLAN network (e.g., flannel), the daemon maintains a VxLAN device named flannel.1 per Kubernetes Node (this is the VNID) and maintains routing for this network. When cross-node traffic occurs, the local node maintains the remote VxLAN device's MAC address to know the destination, encapsulates the packet, and sends it. The remote VxLAN device flannel.1 decapsulates to obtain the real destination.
View the Forwarding database list:
$ bridge fdb
26:5e:87:90:91:fc dev flannel.1 dst 10.0.0.3 self permanentNotes: VxLAN uses port 4789; Wireshark analyzes by port. Flannel's default Linux port is 8472, so captures show only a UDP packet.
The architecture shows that a tunnel is an abstract concept, not a real end-to-end tunnel. It encapsulates packets into another packet, transports via physical devices, and decapsulates via the same device (network tunnel) to achieve network overlay.
weave vxlan
weavealso uses VxLAN for packet encapsulation, called fastdp (fast data path) in Weave. Unlike calico and flannel, it uses the Linux kernel openvswitch datapath module and encrypts network traffic.
Notes: fastdp works on Linux kernel 3.12+. On older kernels (e.g., CentOS 7), Weave runs in user space, called sleeve mode .
Reference
https://github.com/flannel-io/flannel/blob/master/Documentation/backends.md#host-gw https://projectcalico.docs.tigera.io/networking/bgp https://www.weave.works/docs/net/latest/concepts/router-encapsulation/ https://github.com/k8snetworkplumbingwg/sriov-network-device-plugin https://github.com/nokia/danmSigned-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Linux Tech Enthusiast
Focused on sharing practical Linux technology content, covering Linux fundamentals, applications, tools, as well as databases, operating systems, network security, and other technical knowledge.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
