Once a local AI setup grows beyond one machine — an RTX PC plus a unified memory AI node, a NAS holding a model library, or several clients sharing one server — the question of 10GbE comes up quickly. The common assumption is that a faster network makes local AI faster.
That assumption mixes up very different kinds of traffic. Sending prompts and receiving generated text behaves nothing like copying a large model file or editing project assets on shared storage. This page separates them and covers API traffic, model staging, NAS storage, creative assets and multi-client setups.
The goal is to spend on the network only where the network is actually the bottleneck.
The 30-Second Answer
Token generation speed itself is normally not improved by 10GbE.
10GbE matters mainly for model staging, large asset transfers, shared storage and multi-client workflows.
2.5GbE is often sufficient when models live on local NVMe and the network mainly carries API traffic.
The slowest component in the storage/network chain determines actual transfer performance.
In practice: check your storage before upgrading the network. A faster link only helps when the drives at both ends can keep up.
What Actually Crosses the Network
The connection guide separates API requests from model files. For a bandwidth decision, it helps to go one level further:
| Traffic | Typical size | How often | What limits it |
|---|---|---|---|
| API requests and text responses | Small | Constantly during use | Inference speed on the node; connection stability |
| Multimodal payloads (images in prompts, document uploads for retrieval) | Small to medium | Depends on the workflow | Usually inference speed; bandwidth only with large batches |
| Model staging (copying weights to the node) | Large to very large | When adding or swapping models | Network and storage throughput |
| Creative assets and generated output | Medium to very large | Depends on the project | Network and storage throughput |
| Shared storage used live (working directly from a NAS) | Continuous reads and writes | Throughout the session | Sustained network, NAS and drive performance |
| Several clients at once | Sum of all of the above | Overlapping | Links to the shared server or NAS |
Only the bottom four rows are bandwidth problems. Live video streams between real-time systems are a separate, bandwidth-heavy case covered in the TouchDesigner guide. How AI and real-time rendering are split across machines is covered in RTX + Unified Memory for Real-Time Creative Work.
Why 10GbE Does Not Speed Up Token Generation
When an application on the RTX PC sends a prompt, the AI node does the work. Once the model is resident in the node’s memory, generation speed is limited by that node’s memory bandwidth and compute. The network only carries the request in and the text out.
Ordinary text-generation traffic is usually tiny compared with even 1GbE capacity once the model is resident on the AI node.
For this traffic, what matters is a stable wired connection, not a faster one. A higher link speed does not make the model think faster.
Payloads do grow when prompts include images, when pipelines send large batches, when documents are uploaded for retrieval, or when many clients share one node. This can become noticeable, but it is usually still modest compared with copying model files.
Splitting a single model across several machines is a different workload in which the interconnect matters much more. It is outside the scope of the two-PC setup described in RTX PC + Unified Memory AI Node, where each model runs entirely on one machine.
1GbE, 2.5GbE, 5GbE and 10GbE: Line Rate vs Real Throughput
Ethernet speeds are quoted in bits per second; file sizes are quoted in bytes. Dividing the line rate by eight gives a theoretical ceiling; real file transfers typically stay below it because of protocol and system overhead.
| Link | Nominal line rate | Theoretical maximum (line rate ÷ 8) | Illustrative effective throughput used for calculation |
|---|---|---|---|
| 1GbE | 1 Gbit/s | 125 MB/s | 110 MB/s |
| 2.5GbE | 2.5 Gbit/s | About 312 MB/s | 280 MB/s |
| 5GbE | 5 Gbit/s | 625 MB/s | — (less common; not used in the examples) |
| 10GbE | 10 Gbit/s | 1,250 MB/s | 1,000 MB/s |
Illustrative effective throughput used for calculation. Real results depend on network adapters, operating system, file-sharing protocol (SMB/NFS), NAS processor, storage, cabling and switch.
Practical requirements
- 2.5GbE usually works over existing Cat5e or Cat6 cable at typical home and studio distances.
- 10GbE: Cat6A is recommended for new cable runs; Cat6 may work over short distances. 10GBASE-T adapters and switches can run warmer and draw more power than slower ports.
- Every hop counts. A link negotiates to the slowest port on the path: one 1GbE port on a switch or adapter limits the whole connection.
- Jumbo frames are optional. Leave them at the default unless every device on the path is configured consistently.
Cabling, topology and addressing are covered in How to Connect an RTX PC and an AI Node.
Model Staging: Example Transfer Times
Model staging — copying weights onto the node before use — is where link speed becomes visible. The table uses file-size classes rather than parameter counts, because the size of a given model depends on its format and quantization.
| File size | Example | 1GbE (110 MB/s) | 2.5GbE (280 MB/s) | 10GbE (1,000 MB/s) |
|---|---|---|---|---|
| 5 GB | Small quantized model | About 45 s | About 18 s | About 5 s |
| 20 GB | Mid-size quantized model | About 3 min | About 1 min 10 s | About 20 s |
| 50 GB | Large quantized model | About 7 min 30 s | About 3 min | About 50 s |
| 100 GB | Very large model or several checkpoints | About 15 min | About 6 min | About 1 min 40 s |
| 250 GB | Model library or asset archive | About 38 min | About 15 min | About 4 min 10 s |
Example transfer times assuming 110 MB/s, 280 MB/s and 1,000 MB/s effective throughput. Actual results vary by storage, protocol and hardware. Sizes use decimal units (1 GB = 1,000 MB).
How to read the table
An occasional copy is a short wait at any of these speeds. The difference starts to matter when large files move every day: swapping big models repeatedly, restoring archives, or pulling project assets between machines.
The Slowest Link Sets the Speed
A file transfer passes through a chain of components. The slowest one sets the pace for the whole transfer:
Source storage → Source adapter → Cable → Switch → Cable → Destination adapter → Destination storage
Protocol, operating system and processor load apply at both ends.
- Hard drives. A single hard drive, or a small hard-drive array, can be slower than a 10GbE link.
- Sustained SSD writes. Some SSDs slow down during long transfers once their internal cache is exhausted.
- NAS processors. Low-power NAS processors can cap throughput, especially when encryption is enabled.
- Protocol and settings. File-sharing protocol, signing and OS settings all affect the result.
A link that negotiates at 10GbE does not guarantee 10GbE-class transfers. In many local workflows, 10GbE can move the bottleneck away from the network; storage, protocol and CPU performance may then become the limiting factors.
NAS for Model Libraries: Copy-Then-Run vs Run-From-Share
The deciding question for a NAS is whether you copy files from it or work directly on it.
Copy-then-run
NAS → copy → node’s local NVMe → load
The NAS is a library. Network speed only affects how long staging takes. 2.5GbE is often comfortable.
Run-from-share
NAS → load directly over the network
Every model load depends on the network and the NAS. 10GbE and fast NAS storage make this practical.
For most setups, keep active models on the node’s local NVMe and use the NAS for the library and backups. Once a model is loaded into memory, generation is unaffected by where the file came from — but each reload after the model is evicted repeats the network cost.
The same logic applies to creative projects: see Storage Is Part of the Workstation for how working, cache and archive storage are usually separated.
Several Machines, One Storage Pool
When a desktop, a laptop and an AI node all read from the same NAS, their traffic adds up on the link to that NAS. This is where 10GbE most often pays for itself.
- Hybrid design. A common approach is a switch with 10GbE ports for the NAS or AI node and 2.5GbE ports for client machines. The shared device gets the headroom; clients keep simpler hardware.
- Link aggregation. Combining several ports helps when many clients connect at once, but a single transfer between two machines usually does not exceed the speed of one link.
Other Connection Options
Wi-Fi, including Wi-Fi 7: wireless performance varies substantially by environment, so Ethernet provides more predictable sustained throughput. Wi-Fi is reasonable for light API use from a laptop, but the AI node itself and any machine that moves large files should be wired.
Thunderbolt networking: it may be useful for short-distance direct connections on supported systems, such as two machines on the same desk. It is not a replacement for a studio-wide Ethernet network.
Which Link Speed Fits Your Setup?
| If you… | Start with |
|---|---|
| Mostly send prompts and receive text, with models on the node’s local NVMe | 1GbE works; 2.5GbE where your hardware already supports it |
| Add or swap large models occasionally | 2.5GbE |
| Swap large models frequently, or keep a shared model library on a NAS | 10GbE to the NAS or AI node |
| Share storage or creative assets across several machines | A 10GbE backbone with a switch; 2.5GbE clients are often fine |
| Have storage that is already slower than your current link | Upgrade storage before the network |
Measure where the time goes before upgrading
If responses feel slow, look at the node’s inference speed first. If copies feel slow, compare network throughput with what the drives at both ends can sustain. Upgrade the component that is actually waiting.
Frequently Asked Questions
Does 10GbE make local LLM responses faster?
Usually not. Once the model is loaded on the AI node, token generation is limited by the node’s memory bandwidth and compute, not by the network. Ordinary text-generation traffic is usually tiny compared with even 1GbE capacity. 10GbE shortens model staging and large file transfers instead.
Is 2.5GbE enough for a two-PC local AI setup?
Often, yes. When active models live on the AI node’s local NVMe storage and the network mainly carries API requests, 2.5GbE leaves comfortable headroom and makes occasional model copies noticeably quicker than gigabit Ethernet.
Should I run models directly from a NAS?
It works, but every model load then depends on the network and the NAS. Most setups copy active models to the node’s local SSD and use the NAS as a library and backup. Running from a share makes more sense with a fast network, fast NAS storage and models that are loaded infrequently.
Why is my 10GbE transfer slower than expected?
Actual speed is set by the slowest component in the chain: source and destination storage, the NAS processor, protocol settings, network adapters, cables and the switch. A link that negotiates at 10GbE does not guarantee 10GbE-class file transfers.
Can Wi-Fi replace Ethernet for an AI node?
For light API use from a laptop it can work. Wireless performance varies substantially by environment, so a wired connection is the better choice for the node itself and for large transfers, where Ethernet provides more predictable sustained throughput.
Official Technical References
Technical references only. These links support the specifications and concepts on this page and are not purchase links.
- IEEE 802.3 Ethernet Working Group: Ethernet standards, including 2.5G, 5G and 10GBASE-T
- Microsoft Learn: SMB file-sharing protocol overview
- Apple Mac Studio: Technical specifications, including Ethernet options
- Ollama: Official FAQ, including model storage location and network access
Next Steps
Put the network to work in a creative setup, or configure the connection itself: