Kubernetes Network Policies: Default-Deny Egress Blocks DNS
Default-deny egress blocks DNS, causing 30-second timeouts on every request.
- Enforcement is done by the CNI plugin (Calico, Cilium), NOT the API server. Flannel ignores policies silently.
- Default behavior: if no policy selects a Pod, ALL traffic is allowed in both directions.
- Once any policy selects a Pod, that direction enters implicit default-deny. Only explicitly whitelisted traffic passes.
- Policies are additive whitelists. There is no deny rule in the standard API. Multiple policies selecting the same Pod are unioned (OR).
- iptables-based CNIs (Calico) scale O(n) with rule count. Performance degrades at 1000+ Pods.
- eBPF-based CNIs (Cilium) scale O(1) with hash maps. Better performance but requires kernel 4.9+.
- Forgetting DNS egress carve-out when applying default-deny egress. Every service discovery call silently times out after 30 seconds.
Imagine your apartment building has no locks on any doors — every tenant can walk into every other apartment freely. Kubernetes without Network Policies is exactly that: every Pod can talk to every other Pod by default. Network Policies are the deadbolts you install. You decide which apartments can knock on which doors, and everyone else gets turned away at the hallway.
Most teams get Kubernetes running, deploy their apps, and move on — never realizing their payment service can freely dial their logging sidecar, which can freely dial their database, which can freely reach the internet. That's not paranoia; that's the default. Kubernetes was designed for rapid connectivity, not zero-trust isolation. The moment you run multiple tenants, compliance workloads, or anything that touches PII or financial data, that open-door model becomes a liability.
Network Policies solve this by letting you express intent in YAML: only Pods with this label may reach my database on port 5432, from this namespace only, and my database can reach nothing outbound except DNS. The CNI plugin — not the Kubernetes API server — enforces those rules in the kernel using iptables, eBPF, or nftables depending on your stack. That distinction matters enormously for debugging and performance.
This is not a syntax reference. It covers how policies are evaluated and merged, how to write airtight ingress and egress rules without accidentally blackholing DNS, how to verify enforcement at the network level rather than trusting your YAML applied cleanly, and the production mistakes that silently leave clusters wide open.
How Network Policy Enforcement Actually Works — The CNI Layer
Here's the thing most tutorials skip: the Kubernetes API server doesn't enforce Network Policies. It just stores them. The actual enforcement happens inside your CNI plugin — Calico, Cilium, Weave, Antrea — which watches the API server for NetworkPolicy objects and translates them into kernel-level firewall rules on each node.
With Calico on older kernels, that means iptables chains per endpoint. With Cilium, it's eBPF programs loaded into the kernel that intercept packets at the socket layer before they ever hit iptables — significantly lower latency and dramatically better observability. With Flannel, enforcement is zero because Flannel doesn't implement Network Policies at all. This is one of the most common production surprises: a team applies policies and believes they're enforced, but their CNI silently ignores them.
Policy evaluation works like a firewall whitelist. If no NetworkPolicy selects a Pod, all traffic is allowed. The moment any policy selects a Pod — via podSelector — that Pod enters an implicit 'default deny' for the traffic directions that policy governs. Multiple policies selecting the same Pod are unioned together: a packet is allowed if it matches any one of them. There's no precedence, no ordering, no 'deny' rule type in the core API. You get whitelisting only, which is both a simplicity win and a constraint you need to design around.
- Flannel: Provides networking only. No NetworkPolicy enforcement. Zero.
- Calico: Full NetworkPolicy support via iptables or eBPF (with Calico CNI).
- Cilium: Full NetworkPolicy support via eBPF. Extended CRDs for L7 policies.
- Weave: NetworkPolicy support but less performant than Calico/Cilium.
- Antrea: VMware's CNI with full NetworkPolicy support and traceflow debugging.
Writing Precise Ingress and Egress Rules — With the DNS Trap Explained
Once you've applied default-deny, you need to surgically re-open only the traffic paths your application legitimately needs. Ingress rules control what can reach your Pod. Egress rules control what your Pod can reach. Both use the same selector primitives: podSelector, namespaceSelector, and ipBlock, which you can combine with AND logic inside a single from/to entry, or use OR logic across multiple entries.
The subtlety that burns everyone: a from entry with both podSelector AND namespaceSelector means the source must match BOTH selectors simultaneously — it's an AND. Two separate from entries each with their own selector is an OR. The indentation in YAML is load-bearing here. Get it wrong and you either over-permit or under-permit with no error from the API server.
The DNS trap is equally nasty. When you lock down egress, your Pods immediately lose DNS resolution because they can no longer reach CoreDNS on port 53 UDP/TCP. Every connection attempt fails not with a 'connection refused' but with a timeout waiting for DNS — which takes 30 seconds to surface. Always add an explicit egress rule for CoreDNS as part of your default-deny rollout, or you'll wonder why your app is broken when your network policy looks correct.
- Same dash entry with podSelector AND namespaceSelector: source must match BOTH (AND).
- Separate dash entries with podSelector OR namespaceSelector: source can match EITHER (OR).
- No from/to clause under a governed policyType: deny all for that direction.
- Empty from/to clause (from: []): also deny all — same as omitting the clause.
- ipBlock can be combined with podSelector/namespaceSelector in the same entry (AND).
Verifying Real Enforcement and Debugging Policy Failures in Production
Applying a NetworkPolicy and assuming it works is a mistake you only make once in production. The API server accepts any syntactically valid policy regardless of whether your CNI supports it. You need to verify enforcement at the traffic level, not the YAML level.
The gold-standard test is running a temporary Pod in the source namespace and attempting a connection directly — not through a Service mesh or load balancer that might bypass node-level rules. Use kubectl run with --rm -it to spin up a throwaway Pod, then use curl, nc, or wget to probe the target. A dropped connection times out; a policy-permitted connection either succeeds or returns an application-level error (which is actually what you want to see — it means the packet reached the target).
For Cilium clusters, cilium monitor and the Hubble UI are exceptionally powerful — they show you in real time which policies matched or dropped each flow, with source/destination Pod identity, namespace, and labels. For Calico clusters, calicoctl get networkpolicy and iptables -L -n --line-numbers on the node running your Pod reveal the actual enforced rules. Always test both directions — a policy that allows egress from Pod A to Pod B doesn't automatically allow ingress to Pod B from Pod A unless Pod B also has a matching ingress rule.
- Timeout (after 3-30s): Packet was dropped by the CNI. NetworkPolicy is enforcing correctly.
- Connection refused (immediate): Packet reached the target process. NetworkPolicy is NOT blocking this path.
- HTTP 200: Packet reached the application and got a valid response. Policy allows this traffic.
- HTTP 5xx: Packet reached the application but the app returned an error. Policy allows, app has issues.
- DNS timeout (30s): UDP 53 to CoreDNS is blocked. Check egress rules for DNS carve-out.
Production Patterns: Namespace Isolation, Monitoring Carve-outs and Label Hygiene
In a real multi-tenant cluster, you can't write policies Pod-by-Pod. You need namespace-scoped baselines combined with additive per-workload rules. The pattern that works at scale is: one default-deny policy per namespace applied by your CD pipeline at namespace creation, then application-specific policies delivered alongside each Helm chart or Kustomize overlay.
Monitoring is the most common carve-out needed. Prometheus needs to scrape metrics from every namespace, but you don't want to globally allow all ingress. The clean solution is a namespace label like monitoring.io/allow-scrape: 'true' and a policy in each target namespace that allows ingress from the monitoring namespace on port 9090 or whatever your metrics port is. This keeps control local to the target namespace.
Label hygiene is non-negotiable. Network Policies inherit whatever labels your Pods have — if a developer changes a label during a refactor, the policy selector silently stops matching and the Pod falls back to default-deny behavior with no warning event. Use immutable labels like app: payment-api for security selectors and mutable labels like version: v2 only for routing. Audit your selectors in CI with kubectl get pods -l app=api-server -n payments and fail the pipeline if the expected count is zero.
- Security labels (app, tier, team) should be immutable. Enforce with admission webhooks.
- Routing labels (version, canary, blue-green) should NOT be used in NetworkPolicy selectors.
- CI check: fail the pipeline if
kubectl get pods -l app=<name>returns zero Pods. - Namespace labels (kubernetes.io/metadata.name) are auto-applied in Kubernetes 1.21+. Use them for namespaceSelector.
- Adopt a naming convention: all NetworkPolicy names should include the namespace and workload they govern.
Network Policy Performance: iptables vs eBPF at Scale
The CNI enforcement mechanism directly impacts network latency and control plane load. Understanding the performance characteristics of your CNI is critical for capacity planning and troubleshooting latency issues that appear only at scale.
- iptables (Calico default): Sequential rule matching. Degrades at 1000+ Pods per node.
- eBPF (Cilium, Calico with eBPF dataplane): Hash map lookups. Scales linearly.
- iptables rule churn: Every policy change triggers iptables-restore on all nodes. Brief packet drops possible during restore.
- eBPF program updates: Atomic program replacement. No packet drops during policy updates.
- Kernel requirement: eBPF requires kernel 4.9+ minimum. Full features require 5.10+.
cilium_datapath_conntrack_gc_entries and iptables_restore_duration_seconds to detect enforcement bottlenecks.Default-Deny Egress Without DNS Carve-Out: Cluster-Wide Service Discovery Failure
- Default-deny egress blocks DNS by default. Always add a carve-out for CoreDNS on UDP and TCP port 53.
- DNS failure manifests as 30-second timeouts, not immediate errors. This makes it look like a latency problem, not a connectivity problem.
- Readiness probes that use localhost or IP addresses pass even when DNS is broken. Use DNS-based probes to catch this.
- Test default-deny egress in staging with a curl-based smoke test before applying to production.
- CI validation of NetworkPolicy completeness prevents this class of incident entirely.
Key takeaways
Interview Questions on This Topic
Frequently Asked Questions
That's Kubernetes. Mark it forged?
5 min read · try the examples if you haven't