MR2 Solutions
Uncategorized

Sd Wan Design

mr2solutions 16 min read

A CIO with 120 retail branches can see the deadline approaching: an MPLS renewal in 14 months, rising SaaS usage, and business applications that no longer need to pass through the data center. The immediate question sounds technical, “Which SD-WAN platform should we buy?” The better question is, “What network behavior does the business need, and what level of cost, resilience, security, and operational effort can we support?”

That distinction matters. SD-WAN design determines whether broadband becomes a useful second path or an unreliable shortcut, whether cloud traffic takes a sensible route, and whether a new acquisition can be integrated without rebuilding the WAN site by site. It also determines how much telemetry the platform consumes, how quickly teams can detect a failed path, and how much migration risk the organization accepts.

The market context makes this decision more consequential. Industry coverage reported that 87–90% of enterprises had deployed or were actively deploying SD-WAN by 2024, while North America represented about 40–50% of global SD-WAN and SASE spending (Telecom Review Americas). The architecture has moved beyond replacing expensive circuits. Modern designs commonly combine centralized policy, application-aware routing, multiple transports, and security services delivered through a SASE model.

Why SD-WAN Design Decisions Matter Now

A retail CIO facing an MPLS renewal must balance competing needs before selecting a platform. Stores may rely on cloud point-of-sale services, voice traffic may need a direct branch-to-branch path, and security policy may require web traffic inspection. Finance wants predictable operating costs, procurement wants negotiating power, and infrastructure teams need a control plane they can troubleshoot under pressure.

Those requirements produce real design trade-offs. A centralized hub can simplify inspection and compliance, yet add latency for cloud applications. Direct internet access can improve SaaS performance, while expanding the security boundary at each branch. Dual transports increase resilience, but also bring circuit management, monitoring, and recurring costs. No topology wins in every environment.

The market context makes SD-WAN an architectural decision rather than a niche overlay. One industry estimate placed global SD-WAN spending at USD 11.61 billion in 2026, projecting USD 28.32 billion by 2031 and a 19.53% CAGR from 2026 to 2031 (Telecom Review Americas). The figures do not select an architecture, but they explain why organizations are treating WAN design as a strategic platform decision.

Telemetry deserves the same attention as bandwidth and hardware. Application probes, path measurements, logs, and security events consume capacity and create operational data that teams must retain, correlate, and act on. A design that measures everything without defining response thresholds can raise costs without improving incident handling.

Practical rule: Define required traffic behavior first. Select the vendor only after documenting business, application, transport, security, telemetry, and migration requirements.

SASE convergence adds another constraint. Gartner was cited as expecting 60% of new SD-WAN purchases to be part of a single-vendor SASE offer by 2026, compared with 15% in 2022. That shift affects architecture, contracts, skills, and exit planning. Security services should be evaluated with routing and operations from the start, while migration plans should protect existing traffic during phased cutovers.

Gathering Business and Technical Requirements

Start with the business outcome and work toward the architecture. The Cisco SD-WAN Design Guide recommends moving from business goals through application behavior, traffic patterns, transports, site standards, and high-level design constraints (Cisco SD-WAN Design Guide). Following that sequence stops a vendor's default topology from becoming an accidental requirement.

Establish the operating objectives

Ask each stakeholder what must improve, how the change will be measured, and what trade-offs the organization accepts. Useful prompts include:

  • Business leadership: Which customer-facing services must remain usable during a carrier outage?
  • Application owners: Which applications are sensitive to latency, jitter, loss, or asymmetric routing?
  • Security: Which traffic requires inspection locally, through a cloud security service, or at a regional hub?
  • Finance: Is the priority lower recurring transport expense, predictable spending, faster site activation, or less hardware refresh pressure?
  • Operations: Which team owns policy, incident response, certificate lifecycle, and provider escalation?
  • Mergers and acquisitions: How quickly must a newly acquired site connect without adopting the existing routing and addressing model?

Record each answer as a design input, with an owner and a way to verify it. “Improve cloud performance” remains too vague until it names the affected applications, user groups, traffic path, and acceptable service behavior. The same discipline applies to resilience: define which services can tolerate degradation, which require local survivability, and how long an outage can last before the business impact becomes unacceptable.

Telemetry belongs in this requirements phase. Application probes, path measurements, logs, and security events consume bandwidth and generate data that teams must store, correlate, and act on. Set collection scope, retention, alert thresholds, and response ownership before deployment. Measuring every flow can increase operating cost without improving incident handling.

Inventory every site and application

The discovery inventory should include link type and capacity, provider, physical diversity, current CPE, LAN handoffs, routing protocols, addressing, power, rack space, and local support capability. Record steady-state and peak utilization separately. A branch with modest daily traffic may still need different capacity or policy when backups, software distribution, or video meetings create predictable bursts.

Application discovery needs the same level of detail. Classify traffic by business criticality and document expected latency, loss, jitter, bandwidth, inspection, and availability requirements. Identify cloud-bound traffic that currently hairpins through a data center. Capture SaaS dependencies, private applications, voice, video, virtual desktops, IoT, guest access, and management traffic.

Also record migration constraints. Note sites that cannot accept an outage, equipment that must remain in service, routing or addressing dependencies, and teams available for local cutover support. A technically sound target design can still fail if the transition path is unsafe or requires skills that the operating team does not have.

Create traceability before topology

A simple worksheet connects every important design choice to a requirement and an owner.

Requirement Category Discovery Question Design Input Produced
Business continuity Which services must operate during a primary-link outage? Required failover behavior and local survivability
Applications Which workloads are sensitive to latency, jitter, or loss? Application classes and routing thresholds
Traffic flow Which destinations are cloud, branch, data center, or internet bound? Traffic matrix and breakout model
Transport Are circuits physically diverse, and what providers are available? Underlay candidates and path diversity
Security Which segments require inspection or isolation? VRF, firewall, and SASE policy requirements
Operations Who approves, deploys, monitors, and changes policy? RACI ownership and runbook inputs
Growth How will sites, users, applications, and acquisitions change? Capacity model and onboarding standards

This process exposes missing data early. If an application owner cannot provide tolerances, use controlled testing and baseline observation instead of accepting a vendor default. Design confidence comes from traceability, measurable acceptance criteria, and a migration plan, not from the feature count in a proposal.

Choosing the Right Topology and Transport Mix

Topology should follow the traffic matrix. A hub-and-spoke overlay remains sensible when branches need centralized internet breakout, shared security inspection, or controlled access to data-center services. It's easier to reason about and often easier to govern, but it can create a concentration point for cloud traffic and branch-to-branch communication.

Full mesh removes the central transit step. That can suit branch-to-branch voice, video, or interactive applications where an extra hop affects user experience. The price is greater policy complexity, more tunnel relationships to monitor, and more difficult segmentation. Full mesh is justified by traffic behavior, not by the fact that the platform supports it.

Partial mesh is often the practical compromise. Regional hubs, selected branch-to-branch paths, and cloud on-ramps can keep critical traffic direct while retaining centralized controls for destinations that need inspection or compliance. It also supports a gradual migration, because teams can introduce optimized paths where measurements show a real benefit.

A chart comparing hub-and-spoke, full-mesh, and partial-mesh network topologies for optimized SD-WAN traffic management and design.

Match transports to failure modes

MPLS plus broadband can provide a controlled private path and a flexible public path. Dual broadband may lower recurring cost, but only if the circuits use diverse last-mile infrastructure and upstream providers. 5G fixed wireless can help temporary sites, rural branches, or locations where wired diversity is unavailable, but radio conditions and data policies need validation under actual load.

Budget each link for normal traffic, failover traffic, overhead, and expected growth. Don't count a backup circuit as resilient if it can't carry the applications that matter during an outage. Underlay diversity should be verified physically, not inferred from different service names.

For teams evaluating the edge layer alongside broader sales and business-development responsibilities, Hire SDR is a separate resource worth keeping distinct from network architecture. The network decision itself should remain grounded in the documented traffic matrix, not in the presentation style of a vendor demonstration.

A useful technical reference for edge placement and service behavior is the SD-WAN edge overview. In practice, use these rules:

  • Choose hub-and-spoke when centralized breakout, inspection, or regulatory control dominates.
  • Choose full mesh only when measured branch-to-branch performance justifies the operational burden.
  • Choose partial mesh when regional access, cloud on-ramps, and selective direct paths must coexist.
  • Use dual transports when the business impact of an outage exceeds the recurring and operational cost of another circuit.

Designing QoS, Path Selection, and SLA Policies

QoS should express business priorities in rules the edge can enforce. A practical starting model uses eight classes: voice, video, transactional, bulk, interactive, management, default, and scavenger. Map identified applications to DSCP markings, then apply those markings consistently across LAN segments, overlays, and provider handoffs.

The class names matter less than their boundaries. Voice must not absorb unrestricted conferencing traffic, while bulk transfers should yield to transactional flows during failover. Test application identification with encrypted traffic and real workflows. Deep packet inspection that performs well in a lab can lose accuracy when applications change protocols or rely on distributed cloud endpoints.

Treat path selection as an SLA decision

BFD telemetry measures packet loss, latency, and jitter, then uses rolling averages to determine whether a path remains within SLA. As described in the Cisco SD-WAN Design Guide, the default behavior uses a 10-minute poll interval and evaluates six averages before removing a path from SLA. The resulting 60-minute smoothing window limits reactions to brief spikes, but may delay a path change for burst-sensitive applications.

Configure probe behavior according to business impact. Aggressive probes detect failures sooner, but they also create monitoring traffic and consume device processing. A primary tunnel may warrant faster detection than a lightly used backup. Test both settings during congestion, then document the thresholds and expected failover behavior.

Application Class Latency (ms) Jitter (ms) Loss (%) BFD Interval / Multiplier
Voice 150 30 1 500 ms / 3
VDI 150 30 2.5 500 ms / 3
Transactional 150 30 2.5 500 ms / 3
Video 150 30 2.5 1000 ms / 5
Bulk 150 30 2.5 1000 ms / 5

Treat these values as targets from the stated policy model, not universal provider guarantees. Keep queue depth within the agreed interactive tolerance under load, shape voice bursts to the policy rate, and remark non-conforming flows instead of allowing classification to drift.

Path selection also affects the operating model. More probes improve visibility but increase telemetry overhead, while direct cloud paths can reduce latency and complicate inspection as the design converges with SASE. During migration, validate SLA thresholds against both legacy VPN behavior and the new overlay. The SD-WAN versus VPN comparison offers useful architectural context for that trade-off.

Building Redundancy and High Availability

A branch can lose its circuit, edge device, controller reachability, or power. Each failure requires a different recovery mechanism. A second transport cannot replace a redundant edge, and a clustered controller cannot keep a dead circuit forwarding traffic. Design each layer against its own failure domain, then test the combined behavior.

Field results also show why resilience starts before deployment. One report found that only 38% of IT professionals described deployments as fully successful, with inadequate planning and limited internal skills among the major failure drivers, as noted in the Fortinet WAN Transformation Report. The rollout plan, operating skills, and rollback path belong in the availability design.

Compare the options honestly

Design Option Achievable Uptime Incremental Cost Operational Complexity
Single transport Dependent on one provider path Lowest Lowest
Single transport with LTE or 5G failover Improves recovery for selected outages Moderate Moderate
Dual active transports Supports continuous path use and failover Higher Higher
Redundant edge pair with dual transports Protects both device and link failure Highest Highest

Avoid quoting uptime or cost multipliers without a contract and provider-specific model. Compare outage impact, circuit diversity, application criticality, recovery objectives, and internal support capacity. A small office may need one edge with a cellular escape path. A distribution center or clinical site may justify redundant edges and active transports, even though those choices increase operational overhead.

For a plain-language explanation outside enterprise WANs, redundancy for RV travelers provides a useful analogy. The rule applies to branch design as well: a backup that shares the primary's power, carrier, conduit, or upstream device is not independent redundancy.

Protect the control plane

A single orchestrator can create a control-plane gap when a large branch estate depends on one logical service. Use clustered controllers across separate availability zones where the platform supports it. Confirm what branches do during an orchestrator outage. Local survivability should preserve forwarding, required routes, and defined security behavior without waiting for a live policy transaction.

Edge high availability also requires a forwarding model. Active/active tunnel termination uses capacity efficiently, but it demands careful flow handling and can expose asymmetric paths. Active/standby simplifies state management, while fast routing convergence limits disruption. Where stateful inspection cannot tolerate asymmetry, use per-flow stickiness or front-haul load-balancer hashing.

Telemetry has a cost. More health probes improve failure visibility but consume bandwidth and device processing, especially across many branches. Set probe frequency according to application impact, then test detection and recovery while links are congested. Direct cloud paths may reduce latency while making inspection and SASE integration more complex, so record who owns policy enforcement on each path.

For brownfield sites, compare early contract termination, temporary dual transport, and staged migration. The break-even point depends on termination penalties, remaining contract term, circuit prices, migration labor, outage exposure, and the cost of operating both designs during stabilization. Calculate those inputs by site class, not as one enterprise-wide average.

Embedding Security and SASE into the Design

Encrypted overlay tunnels protect data in motion. They don't automatically inspect east-west traffic, authenticate every endpoint, or stop a compromised IoT sensor from reaching a flat branch subnet. The statement “SD-WAN is secure by design” is therefore incomplete.

Start by mapping actual entry and movement paths. A rogue device may enter through a guest VLAN. DNS tunneling may use an otherwise permitted underlay. A compromised camera or building controller may attempt lateral movement toward clinical, operational, or corporate systems. Each path needs a control, an owner, and a test.

Build segmentation into the overlay

Use VRF-aware overlays or equivalent segmentation to separate users, voice, guest access, IoT, management, and operational technology. Apply least-privilege rules between segments, not only between branches. Add 802.1X or another device-authentication and posture process at the access layer, so the SD-WAN edge doesn't become the first place where an unknown endpoint is noticed.

Certificate lifecycle deserves the same rigor as routing policy. Define enrollment, renewal, revocation, device replacement, and emergency recovery procedures. Where supported, use dual-issuer redundancy, OCSP stapling, and a 30-day rotation discipline enforced by policy. The exact implementation depends on the platform and certificate authority, but the operational requirement is constant: distributed edges must not rely on manual renewal by local technicians.

A five-step diagram illustrating the process of embedding security and SASE into network design.

Converge in controlled stages

SASE convergence works better as a sequence than as a simultaneous replacement of every security function. A practical path begins with a cloud-delivered secure web gateway for web egress, adds a cloud access security broker for SaaS governance, and then evaluates zero-trust network access as a replacement for remote-access VPN. A 12 to 18 month ZTNA transition window can provide time to map users, devices, applications, and exceptions before the old access model is removed.

The design must also account for platform concentration. Industry analysis describes SD-WAN moving toward secure access integration, autonomous operations, and AI-based traffic optimization, while market coverage highlights risks including controller concentration, internet-reachable edge devices, and misconfiguration that can create flat branch-to-OT access (Research and Markets SD-WAN Analysis). Security teams should review controller privileges, administrative access, patch latency, segmentation defaults, and failure behavior when the underlay is unstable.

Organizations building specialized technical teams may also review resources related to recruiting top MFE graduates, but hiring alone won't solve architectural risk. The design needs enforceable policy, tested recovery procedures, and clear accountability.

Validating, Rolling Out, and Governing the Design

A branch rollout can pass design review and still create outages if validation is rushed or stabilization is compressed. Parallel migrations increase operational exposure, while a site-by-site sequence gives engineers time to observe application behavior, correct policy errors, and update the runbook before the next wave. Rollout speed matters only when the operations team can absorb the risk.

A four-phase diagram outlining the process for validating, rolling out, and governing network design projects.

Use four gated phases

Phase 1, lab validation. Reproduce representative transports, encryption, security inspection, application mixes, and failure conditions. Use synthetic traffic and packet captures to verify application identification, DSCP preservation, queue behavior, path selection, failover, and certificate rotation. Test the control plane separately from data-plane forwarding.

Phase 2, controlled pilot. Select branches with different circuit types, user profiles, and local constraints. Include a site that exposes likely production problems, rather than choosing only technically clean locations.

Phase 3, regional rollout. Gate each wave on application performance, failover under load, inspection accuracy, SLA breach behavior, help-desk volume, and policy consistency. Version the migration runbook, and define rollback triggers before each change window.

Phase 4, governance and optimization. Schedule design reviews, change control, capacity planning, and telemetry-based policy updates. Assign a RACI across network design, operations, security, providers, and application owners. The runbook should address incident triage, path-policy drift, certificate failure, provider escalation, device replacement, and branch onboarding.

Budget the invisible operating costs

The academic SD-WAN survey recommends keeping data-plane monitoring overhead below 5% of total capacity (SD-WAN Survey Postprint). Use that constraint when setting probe placement, sampling, retention, and alert thresholds. Extra visibility can consume bandwidth and CPU, while excessive alerts train operators to ignore warnings.

Governance also includes telemetry retention, vCPE lifecycle management, software and certificate upgrades, controller access, and provider performance. Managed operations may fit the operating model, and managed services operations offers context for assigning monitoring and support responsibilities.

A software-defined WAN changes through policy, code, templates, and measured feedback. Treating it as static circuits creates drift. Govern it as a software asset, with ownership and review gates, so the business can change paths, add sites, adopt cloud security, and retire legacy connectivity without losing control.

MR2 Solutions helps organizations evaluate, design, procure, implement, and govern vendor-neutral SD-WAN and enterprise networking solutions around business requirements. Visit MR2 Solutions to turn traffic matrices, transport choices, security models, and rollout risks into an actionable WAN design and delivery plan.

Join the conversation

Your email address will not be published. Required fields are marked *

Ready to make smarter technology decisions?

Talk to a vendor-neutral advisor about your next technology initiative.

Schedule a Consultation