Skip to content
Tech Interview Prep home
Technical interview guide

Cloud Networking Fundamentals

VPCs, subnets, and security groups — the building blocks every other cloud topic assumes.

Read
22 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: Vendor-neutral fundamentals with AWS VPC terminology and official references accessed 2026-08-30..

Overview

Curated: · Written: · Reviewed:

Key takeaways

  • A VPC is a routing and isolation boundary; a subnet is an address range placed in one failure domain.
  • A subnet is public because its effective route table can send traffic to an internet gateway—not because of its name.
  • A public IPv4 route is necessary but not sufficient: the resource also needs a public address and filtering that permits the flow.
  • NAT enables private IPv4 workloads to initiate outbound connections. It does not make them directly reachable from the internet.
  • Routes decide where packets go; stateful security groups and stateless network ACLs decide whether packets may pass.
  • CIDR planning is an architectural decision. Overlap becomes expensive when networks later need peering, VPN, or transit connectivity.
  • Design DNS, endpoints, egress, flow logging, and failure domains alongside addresses and firewalls—not after deployment.

1. Start with traffic flows, not a subnet diagram

Interviewers rarely open with "explain a VPC." They open with a scenario: "Your instance can't reach an external API—walk me through debugging it," or "Design the network for a three-tier application." The candidates who do well start from flows, not from a diagram of boxes.

Before selecting CIDR blocks, list the required flows: user to edge, edge to application, application to data, workload to cloud service, workload to internet, operations to workload, and cloud to on-premises. For each flow record source identity or range, destination, protocol and port, direction, latency expectation, sensitivity, and behavior during dependency failure. That inventory drives routing, filtering, name resolution, egress controls, and observability.

A VPC is a logically isolated virtual network. Its address ranges contain subnets, and its router evaluates the route table associated with each subnet. A subnet normally belongs to one availability zone even when the VPC spans a region. Redundancy therefore requires equivalent subnets and capacity across multiple zones; putting two instances in one subnet does not protect against a zone-level failure.

A weak answer here sounds like a vocabulary recital—"a VPC is a virtual network in the cloud"—with no mention of zones, routing, or what happens when a dependency dies. If you catch yourself defining terms instead of describing flows, restart from the traffic.

2. The layered model as a diagnostic map

You don't need to recite all seven OSI layers, but you do need the working map interviewers expect when they ask "where would you look for this failure?":

  • Link/network reachability — does a route exist in both directions? Is the address what you think it is?
  • Transport — is the listener bound to the right address and port? Is something completing the TCP handshake?
  • Name resolution — did DNS return the address you expected, recently enough?
  • Filtering — do the security group, NACL, and host firewall each permit the flow?
  • Application — TLS, certificates, auth, and the actual request/response.

The value of the map is elimination order. A TLS certificate error is not a routing problem; a connection timeout at SYN is not an application problem. When an interviewer hands you "the service is down," naming the layer you'd rule out first—and why—is the answer they're listening for. A weak answer jumps straight to one layer (usually the firewall) without saying what evidence would confirm or rule it out.

3. Address planning and CIDR arithmetic

CIDR expresses a network prefix such as 10.20.0.0/16. A longer prefix is a smaller, more specific range: 10.20.8.0/24 sits inside that /16. Route lookup uses the most specific matching prefix, so a route for 10.20.8.0/24 wins over a route for 10.20.0.0/16 when both match.

Be ready to do the arithmetic live. A /24 holds 2^(32−24) = 256 addresses; cloud platforms reserve a few (AWS reserves five per subnet: the network address, the VPC router, DNS, the future-reserved address, and the broadcast address), leaving 251 usable. A /16 holds 65,536. Splitting a /16 into /24s gives 2^(24−16) = 256 subnets. If an interviewer asks "how many /26s fit in a /22, and how many hosts each," the answer is 2^(26−22) = 4 subnets of 2^(32−26) = 64 addresses each. Practice this until it's reflexive—it's a common warm-up question.

RFC 1918 reserves 10.0.0.0/8, 172.16.0.0/12, and 192.168.0.0/16 for private IPv4 use. Choosing from those ranges does not make an application secure; it only means the addresses are not globally routed. Avoid overlap with other VPCs, business units, acquired networks, and on-premises ranges. Peering and routed VPN designs generally cannot resolve ambiguous overlapping destinations without translation or renumbering.

Leave room for growth, but do not create one enormous flat failure and trust boundary. Allocate predictable, non-overlapping blocks by environment, region, zone, and tier. Record ownership in an IP address management system. Dual-stack designs also need explicit IPv6 routes and filtering: 0.0.0.0/0 says nothing about IPv6, whose default route is ::/0.

4. Public and private are routing properties

In the AWS model, a subnet is public when its route table has a route to an attached internet gateway. For IPv4, an instance also needs a public IPv4 address for internet communication. A route alone does not assign one, and a public address alone does not create a route. Security controls must independently allow the flow.

A private subnet lacks a direct internet-gateway route. Private IPv4 workloads that need outbound updates or third-party APIs can route through a NAT gateway in a public subnet. The NAT device translates the source and allows return traffic for connections initiated from inside; it does not accept arbitrary new inbound connections to those private workloads. NAT is an availability, throughput, logging, and cost dependency. Deploy and route it deliberately per failure domain rather than silently sending every zone through one node.

When workloads access supported cloud services, private service endpoints can avoid public egress and reduce the permissions and destinations exposed through NAT. They do not automatically solve application-layer authorization: endpoint policy, service policy, workload identity, and resource policy still matter.

The classic interview probe here: "An instance has a public IP but can't reach the internet—why?" There are at least three independent answers (no IGW route, no public address, filtering blocks it), and the strong candidate enumerates them instead of guessing one. The reverse probe—"why can't the internet reach my instance in a public subnet?"—usually turns on the security group, since the route and address are present.

5. Routing and connectivity

Each route has a destination prefix and a target such as the local VPC router, internet gateway, NAT gateway, peering connection, transit hub, VPN, or network interface. Debug routing in both directions. A forward path without a return route produces timeouts that often look like firewall failures.

Peering is direct connectivity between two networks and usually is not transitive: if A peers with B and B peers with C, A does not automatically reach C through B. Interviewers love this one as a follow-up to any peering answer—"and what if C needs to reach A?" The weak answer assumes transitivity or hand-waves "transit routing" without naming the shared-dependency trade-off. A transit hub centralizes connectivity and policy for many networks but introduces a shared dependency and route-domain design problem. Site-to-site VPN gives encrypted connectivity over the internet; dedicated private circuits can provide more predictable capacity, but neither eliminates the need for redundant tunnels, dynamic routing, monitoring, and failover tests.

6. Filtering: stateful and stateless layers

A security group is normally attached to a network interface or resource and is stateful: response traffic for an allowed connection is tracked automatically. Prefer rules that reference workload identities or security groups when the platform supports it, rather than permitting a broad changing address range.

A network ACL operates at the subnet boundary and is stateless. Inbound and outbound directions are evaluated separately, so return traffic—including ephemeral client ports—must be permitted explicitly. ACLs can provide a coarse guardrail or emergency deny, but complicated ACL rule sets are easy to misorder and difficult to debug. Neither control replaces host firewalls, workload identity, TLS, or application authorization.

Expect the follow-up: "Why did my connection fail when the NACL clearly allows the return traffic?" The answer is that the response leaves on an ephemeral port (typically 1024–65535), and a stateless outbound rule permitting only the service port drops it. Security groups never hit this because connection tracking admits the reply. If you can explain that asymmetry with the ephemeral-port detail, you've answered the stateful/stateless question at the depth being probed.

7. DNS and service discovery

Applications usually connect by name, so DNS is part of the availability path—and it's the step most often skipped in debugging. Before touching routes or firewalls, confirm what address the client actually resolved. Split-horizon DNS may return private addresses inside the network and public addresses outside it. Private zones must be associated with the intended networks, forwarding rules must cover on-premises names, and resolvers need observable failure behavior. A healthy route to an IP does not help if resolution is wrong, stale, or blocked.

Use stable service names rather than embedding instance addresses. Keep DNS TTLs aligned with failover goals, but remember that clients and intermediate resolvers may cache. Test actual client behavior during endpoint replacement instead of assuming the authoritative TTL guarantees recovery time.

8. Production design and troubleshooting

Minimize internet-exposed resources: commonly only managed edge/load-balancing components belong in public subnets, while application and data tiers remain private. Restrict egress by destination and purpose where practical; unrestricted outbound access turns a compromised workload into an easy exfiltration path.

Enable flow logs with retention and access controls, but understand their granularity and delivery delay. Combine them with load-balancer, DNS, NAT, firewall, and application telemetry. For a failed connection, check in order: name resolution, source and destination addresses, forward route, source filter, destination filter, listener/process, return filter, return route, and any translation or asymmetric path. Test from the real source because effective routes and policies differ by subnet and identity.

Treat network configuration as version-controlled infrastructure. Validate non-overlap, route intent, forbidden public exposure, and broad ingress/egress in CI; deploy incrementally; retain a tested rollback; and run failure exercises for a zone, NAT path, VPN tunnel, resolver, and transit dependency.

When an interviewer asks "how would you design this network," they are usually scoring whether you mention failure domains, egress control, and observability unprompted. A design that only covers subnets and CIDRs reads as junior regardless of how clean the address plan is.

Worked example: a public address is not a public subnet

Instance 10.20.8.25 has 203.0.113.44 assigned. The subnet route table has only the local VPC route.

setupoutbound IPv4 to 8.8.8.8
public address, no 0.0.0.0/0 to internet gatewaytimeout
0.0.0.0/0 to internet gateway, no public addresstimeout
both, security group allows 443SYN then handshake

The name of the subnet is not the control. The route, the address, and the filter each have to be present.