X4NETWORKSStart a conversation →
JUMP TOProblems we solveAI networkingUFM walkthroughCase filesFailure labNetworkSecurityCloudDocumentationOur processBlogContact →
CCIEs · AVAILABLE 24 × 7

The network behind your next big move.

An AI cluster. A new office. A security overhaul. Whatever is next, we design it with your team—and stay for the hard parts.

X4 Networks engineers reviewing a design with a client IT team
Built with you.The design review, the change window, and the call afterward—all with the same engineers.
25 yearsengineering production networks
CCIEsdirectly involved in every engagement
24 × 7reachable when the network is not
WHAT GOOD INFRASTRUCTURE FEELS LIKE

Your people move faster. Your applications stay available. And the network stops being the thing everyone worries about.

X4 Networks engineer working in a datacenter rack
ENGINEERING WITHOUT THE HANDOFF

The people in the room are the people doing the work.

We do not send a polished design over the wall and disappear. We learn the environment with your team, make the choices together, and carry them through the rack, the console, and the change window.

“You should never have to explain your network twice.”
WHEN TEAMS CALL X4

The work that never quite reaches the top of the list.

Usually nothing is on fire. Something is simply slower, riskier, or harder to explain than it should be. We step in beside your team and close it.

01 / PERFORMANCE

The branch that crawls every afternoon.

We trace the traffic, isolate the constraint, and fix the actual bottleneck—not the loudest symptom.

02 / DELIVERY

The project that has been 80% done since spring.

Senior engineering capacity to carry the design through implementation and a clean handback.

03 / VISIBILITY

The diagram nobody trusts anymore.

We rebuild it from the live environment, layer by layer, until the picture matches the racks.

04 / SECURITY

Firewall rules nobody wants to remove.

Policy reviewed with context, exposure closed, and every remaining rule made explainable.

05 / OPERATIONS

The issue nobody can pin down.

Routing, switching, security, and cloud investigated as one system instead of four queues.

06 / ACCOUNTABILITY

Vendors pointing at one another.

One senior engineer owns the thread, gathers the evidence, and keeps it moving to resolution.

LIVE FABRIC MODEL · BALANCED FAT TREE

Every leaf reaches every spine.

Spine count follows port radix and aggregate bandwidth—not leaf count. This two-leaf example uses bundled uplinks to preserve 1:1 capacity in the healthy state.

8 × 400G DOWN / LEAFGPU-FACING CAPACITY · 3.2 Tb/s
4 × 400G TO EACH SPINE8 UPLINKS / LEAF · 3.2 Tb/s TOTAL
1:1 HEALTHY FABRICAGGREGATE UP = AGGREGATE DOWN
S1Quantum-2 · 4×400G / leaf
S2Quantum-2 · 4×400G / leaf
L1Rail 01 · 8×400G up / 8×400G down
L2Rail 02 · 8×400G up / 8×400G down
⚡ GPU 01–08
⚡ GPU 09–16
50%VIA SPINE 1
50%VIA SPINE 2
NON-BLOCKING CHECK · 3.2 Tb/s uplink = 3.2 Tb/s downlink per leaf
Balanced across both spines · 1:1 aggregate capacity
ONE ENGINEERING PRACTICE

From GPU to user.
No gaps between teams.

01

AI fabric

InfiniBand and Spectrum-X engineered so every GPU spends its time computing—not waiting.

Discuss the work →
02

Core network

Datacenter, campus and WAN designed as one system, with no mystery hops.

Discuss the work →
03

Security

Identity, segmentation and policy that protect the business without slowing it down.

Discuss the work →
04

Cloud edge

A clean, observable path between users, workloads, providers and the internet.

Discuss the work →
DOCUMENTATION THAT TELLS THE TRUTH

Your network is not
one drawing.

It is three stories happening at once. We verify each one against the live environment—then make them readable by humans.

X4 / AS-BUILTVERIFIED 08.2026
CORE NETWORK · PHYSICAL
INTERNET2 carriers
FIREWALLHA pair · rack A1
CORE2× 100G DAC
BRANCHEScarrier handoff
USERSCat6A · PP-02
SOURCE / live configuration + cable auditSHEET 1 OF 3
AI NETWORKING · END TO END

The fabric is part of the machine.

GPU performance is not only a server question. The network determines how quickly accelerators exchange data, how well the cluster handles a failed path, and whether expensive compute waits on communication.

NVIDIA INFINIBANDSPECTRUM-XNDR 400GRDMARAIL-OPTIMIZEDADAPTIVE ROUTING
X4 Networks engineers reviewing AI-fabric operations together
Designed in the room. GPU count, workload shape, growth plan, power, cabling, and operations considered together.
01

Cluster discovery

Translate training and inference goals into bandwidth, oversubscription, topology, port, cable, and growth requirements.

02

Fabric architecture

Rail-optimized leaf-spine designs, lossless transport, adaptive routing, isolation boundaries, and out-of-band management.

03

Physical implementation

Rack elevations, port maps, optics, cable plans, labeling, staging, and installation that remains understandable at full density.

04

Validation under load

Confirm link health, path balance, latency, congestion behavior, failover, and end-to-end RDMA performance before the first production job.

05

Operational readiness

As-built documentation, monitoring, baselines, upgrade plans, escalation paths, and runbooks your infrastructure team can own.

NVIDIA UFM · GUIDED OPERATIONS PREVIEW

See the fabric before users feel it.

Click through the views—or play the guided walkthrough—to see how an operations team moves from cluster health to topology, traffic, and a developing congestion event.

FABRIC HEALTHNETWORK MAPTELEMETRYEVENTS + ALARMS
GUIDED UFM OPERATIONS PREVIEW · AI-FABRIC-01

Fabric Dashboard

● FABRIC HEALTHY

Traffic Map SPINE → LEAF → GPU NIC RAIL

SPINE-01
SPINE-02
LEAF-01RAIL 01
LEAF-02RAIL 02
LEAF-03RAIL 03

Inventory ACTIVE

64GPU NODES
18SWITCHES
0DOWN LINKS

Recent Activity LIVE

  • Adaptive route updatedrail-02 · 12 sec ago
  • Daily fabric report readyoperations · 3 min ago
  • All managed elements reachablehealth check · 4 min ago

Network Map SPINE → LEAF → GPU NIC RAIL

QM9700-SP01
QM9700-SP02
LEAF-01RAIL 01 · 400G
LEAF-02RAIL 02 · 400G
LEAF-03RAIL 03 · WATCH
GPU NICsRAIL 01
GPU NICsRAIL 02
GPU NICsRAIL 03

Fabric Throughput TB/S

Top Links by Utilization 30 SEC

Congestion Watch 1 DEVELOPING

  • RAIL-03 / PORT 2792% utilization
  • Adaptive routingredistributing flows
  • Packet discard0 detected

UFM Health NOMINAL

100 / 100

Fabric Health WATCH

98 / 100

Operations READY

  • Topology compareNo unexpected changes
  • Daily reportGenerated 06:00
  • SnapshotAvailable for support
READY · CLICK ANY VIEWConceptual interface · representative data · not a live UFM session
NETWORK VISIBILITY

Documentation your team can actually use.

A useful network drawing answers the next question before someone has to ask it. We map the live environment from physical connections through switching, routing, security policy, and traffic flow.

01Current-state topology verified against live configuration
02Routing policy and traffic flow explained in plain language
03Runbooks built for changes, incidents, and audits
X4 Networks engineer reviewing network monitoring on dual displays
LIVE ENVIRONMENT REVIEW
THE REST OF THE PRACTICE

One team across the infrastructure.

The hard problems rarely respect product boundaries. Our engineers work across the path so an issue does not disappear between specialties.

SECURITY

Zero Trust and firewall policy

Identity-led access, useful segmentation, Palo Alto and Cisco firewall engineering, rule-base remediation, and ongoing support.

Talk through the exposure →
CONNECTIVITY

SD-WAN that earns its keep

Application-aware routing, predictable branch rollouts, cleaner operations, and honest visibility into how the fabric behaves.

Improve the fabric →
CLOUD

Cloud without a blind spot

Infrastructure, private connectivity, backup and replication, and policy carried consistently across providers.

Connect the environments →
ENGINEERING CASE FILES · LAB-NORMALIZED DATA

Show the evidence. Then make the change.

These reference cases expose the complete engineering chain: symptom, topology, counter evidence, design decision, implementation and acceptance result.

TRANSPARENCY NOTE · These are normalized lab case files, not customer claims. Replace the values with approved client data when available.
CASE 01 · AI FABRIC

The GPUs were fast. The collective was not.

64 GPUs · 8 racks · 400G IB
SPINE 1SPINE 2RAIL 1RAIL 2RAIL 4 GPU 01–16GPU 17–32GPU 33–48GPU 49–64
01 · ORIGINAL PROBLEM

Training iterations developed a long tail while average link utilization still looked reasonable.

02 · EVIDENCE

Normalized XmitWait concentrated on two host-facing ports; NCCL p99 aligned with the same windows.

03 · DESIGN DECISION

Restore rail locality, move two HCAs to the correct leaf pair and enable the validated adaptive-routing profile.

04 · IMPLEMENTATION

Recable by rail, verify GUID-to-port mapping, run ibdiagnet, then repeat the collective acceptance suite.

61% → 86%GPU UTILIZATION
42 → 19 msCOLLECTIVE P99
8.4% → .7%NORMALIZED XMITWAIT
0FAILED ACCEPTANCE TESTS
CASE 02 · CORE NETWORK

A redundant design with an eighteen-second interruption.

2 DCs · 24 leaves · dual carriers
CARRIER ACARRIER BCORE ACORE BDATA CENTER 1DATA CENTER 2
01 · ORIGINAL PROBLEM

Carrier maintenance produced an application-visible interruption despite dual circuits and dual cores.

02 · EVIDENCE

Packet capture showed the data plane waiting on recursive next-hop repair after the physical loss.

03 · DESIGN DECISION

Align BFD, routing timers and next-hop tracking with the application loss budget; remove hidden recursion.

04 · IMPLEMENTATION

Stage timers, fail each direction independently, capture loss bursts and document the rollback threshold.

18.6 → 1.2 sCONVERGENCE
2.9% → .03%LOSS DURING TEST
7 → 1MONTHLY INCIDENTS
2 pathsVALIDATED BOTH WAYS
CASE 03 · SPECTRUM-X / RoCE

The pause storm looked like a GPU problem.

32 GPUs · 100G RoCEv2 · shared storage
SPINE 1SPINE 2LEAF 1LEAF 2GPU PODSTORAGECPU FARM
01 · ORIGINAL PROBLEM

GPU utilization collapsed during checkpoint writes; retransmits and pause duration rose together.

02 · EVIDENCE

One optic showed symbol errors while PFC pause frames spread the symptom across otherwise healthy paths.

03 · DESIGN DECISION

Replace the physical fault first, then validate ECN thresholds, PFC headroom and queue-to-priority mapping.

04 · IMPLEMENTATION

Swap cable, clear counters, replay storage load, prove no-drop behavior and record queue telemetry.

58% → 84%GPU UTILIZATION
284 → 0SYMBOL ERRORS / RUN
11 → 2JOB RESTARTS / WEEK
0PACKET DROPS IN ACCEPTANCE
INTERACTIVE AI FABRIC FAILURE LAB

Break the fabric. Follow the evidence.

Trigger a modeled failure and watch the path, counters and CCIE investigation sequence change together.

CONCEPTUAL LAB · Counter names follow NVIDIA UFM semantics. Values are normalized demo data, not live customer telemetry.
X4 LAB · 2 SPINES / 4 RAILS / 32 GPUsFABRIC HEALTHY
SPINE 1SPINE 2RAIL 1RAIL 2RAIL 3RAIL 4 GPU 01–08GPU 09–16GPU 17–24GPU 25–32
0.2%NORMALIZED XMITWAIT
0.1%CONGESTED BANDWIDTH
0SYMBOL ERRORS
0LINK-DOWN EVENTS
92%GPU UTILIZATION
18 msCOLLECTIVE P99
ANNOTATED TECHNICAL EVIDENCE

Not a screenshot. A conclusion with a trail.

Every artifact is paired with the question it answers. Official interface evidence is identified; lab captures are labeled; acceptance criteria stay visible.

Official NVIDIA UFM telemetry view

Official NVIDIA UFM telemetry time-range interface screenshot 01 · Set the time window before correlating a counter spike.02 · Compare the affected port with its rail peers—not the fabric average.
Official NVIDIA documentation image, shown for technical annotation. Product names and interface remain NVIDIA property.

Sanitized lab telemetry capture

OBJECT leaf03 / port17 LINK 400G / active NormalizedXW 0.084 Normalized_CBW 0.068 PortXmitPktsRate 118.4 Mpps SymbolError 0 LinkDowned 0 CORRELATION NCCL p99 +23 ms CONCLUSION credit pressure, not optics
The absence of physical errors narrows the branch. XmitWait is evidence of insufficient credits or arbitration delay; it is not a diagnosis by itself.

Acceptance-test record

TESTTHRESHOLDRESULT
Spine failover< 2.0 s1.2 s · PASS
Collective p99< 25 ms19 ms · PASS
Normalized XmitWait< 1.0%0.7% · PASS
Symbol errors0 / run0 · PASS
Rail imbalance< 8%6% · WATCH
A design is complete only when the pass criteria, test method and rollback boundary are recorded together.

UFM terminology is grounded in the NVIDIA UFM Enterprise telemetry documentation and the supported port-counter reference.

HOW WE WORK

From the first walkthrough to the runbook.

Four deliberate steps leave the environment more reliable—and your team better able to understand and operate it.

01 / DISCOVER

Start with what is real.

Walk the environment, pull live configuration, and listen to the people who run it every day.

02 / DESIGN

Make the tradeoffs visible.

Put the topology, risk, cost, and operational impact on the same page before a change begins.

03 / DELIVER

Carry it through production.

Stage, implement, validate, and stay close through the change window and what comes after.

04 / SUPPORT

Keep the context.

The same senior engineers remain reachable 24 × 7, without making your team start the story again.

X4 Networks engineers working through a network design with a client team
WHO YOU WORK WITH

Experience is most useful when it stays in the room.

Our clients know the engineers by name. The person who learns the environment is the person who designs the change—and the person who picks up when something unexpected happens.

CCIE-led, not CCIE-branded.Senior certification is present in the work itself, not parked in a capabilities deck.
Meet the team behind the change →
WHY X4

CCIE-led.
Straight answers.
No handoffs.

Your CCIE stays involved from the design review through the change window.

We document what is running—not what the old file says should be running.

We stay accountable after launch, with 24 × 7 access when the stakes are high.

WAYS TO WORK WITH X4

Bring us the outcome, not a perfectly written scope.

Some clients need a full design and implementation. Others need one difficult project closed or a senior engineer who already knows the environment when the next issue appears.

DESIGN + BUILD

Own the change end to end.

Architecture, staging, implementation, validation, documentation, and operational handoff with one accountable team.

BEST FOR / AI FABRICS · DATACENTER · MAJOR REFRESH
PROJECT RESCUE

Close the work that stalled.

Independent assessment, a clear recovery plan, and senior hands to move a delayed or high-risk project over the line.

BEST FOR / MIGRATIONS · POLICY CLEANUP · PERFORMANCE
ONGOING SUPPORT

Keep senior context on call.

Planned changes, escalation, upgrades, and 24 × 7 support from engineers who already understand the environment.

BEST FOR / LEAN TEAMS · COMPLEX NETWORKS · COVERAGE
FOLLOW THE TRAFFIC

See the network make a decision.

Choose a workload. The path, policy, and performance target change with it—because not every packet should be treated the same.

GPU NODEsource
FABRIC LEAFlossless queue
SPINEadaptive route
GPU NODEdestination
LOW LATENCY TARGET
RDMA POLICY
01 QUEUE

Large collective operations stay on the lossless fabric while adaptive routing balances congestion in real time.

TWO-MINUTE SELF CHECK

How much of your environment could you explain today?

Answer five quick questions. Nothing is saved or sent; the score stays in your browser.

Our diagrams match what is running now.

We can explain the owner and purpose of every firewall rule.

Branch, datacenter, and cloud are monitored as one environment.

A failed path has been tested—not just designed.

Our oldest open network project has a clear owner and next step.

WHAT’S THE THING THAT KEEPS COMING BACK?

Let’s make it
the last time.

Talk with an engineer →609.683.9000