Creators increasingly connect local language models to TouchDesigner networks, Unreal Engine projects and 3D pipelines — for dialogue, cue generation, procedural parameters or an always-available assistant. When the model and the creative application run on the same GPU, they draw on the same resources at the same time.
This page covers how to divide that work between an RTX creative workstation and a unified memory AI node: what each machine should run, how the software should talk to the node, and where the approach stops working. It is not a per-application hardware guide; for sizing a creative workstation, see Creator PCs in Japan.
The aim is to keep large-model inference from competing with the real-time frame budget.
The 30-Second Answer
Real-time rendering, the viewport, GPU rendering and CUDA-first tools such as most image-generation pipelines.
Large language models, long-context inference and a persistent agent backend that stays loaded all day.
Asynchronous API calls over the LAN, so inference does not have to block the render loop.
The limit to keep in mind: separation protects the frame budget from AI load; it does not make the model faster.
Why Real-Time Work and AI Inference Collide on One GPU
Real-time applications and AI inference can compete for VRAM, compute resources and memory bandwidth on the same GPU.
- VRAM. Textures, render targets and feedback buffers share memory with model weights and context. When combined demand exceeds capacity, something has to give: the model is unloaded, or the scene has to fall back.
- Compute. Inference work and rendering work are scheduled on the same GPU. While a long inference request runs, rendering has less GPU time available.
- Memory bandwidth. Both workloads move large amounts of data through the same memory system.
Depending on the scene and the model, the result can be uneven frame pacing, slower renders, delays while models reload, or out-of-memory errors when combined demand exceeds VRAM. How much of a problem this is varies by project — see VRAM planning for TouchDesigner for what drives memory use on the creative side, and the single-machine ceiling for the general case.
The Real-Time Frame Budget
Real-time work runs against a clock. Each frame has to finish within a fixed interval set by the target frame rate — at 60 frames per second, that interval is about 16.7 milliseconds. Any GPU work that lands in that window competes with rendering for it.
Language-model requests behave very differently from rendering. They usually take far longer than a single frame and arrive at unpredictable moments, driven by an audience, a performer or a script. That mismatch is the main reason to keep large-model inference off the creative GPU rather than trying to fit it between frames.
What Separating the Machines Changes
Changes
- Isolates large-model inference from the creative GPU
- Protects the real-time frame budget from AI load
- Reduces the need to unload and reload large models
- Keeps AI software dependencies apart from the creative environment
Does not change
- Contention caused by the creative scene itself
- How fast the model generates responses
- AI workloads you choose to keep on the RTX side
Adds
- A network dependency between the machines
- A second machine to maintain and update
- Power planning for two systems
In Japan, two high-end machines on household circuits are worth planning for; see High-End PCs on Japan’s 100V Power.
Typical Workload Placement
These are default placements, not rules. Each workload goes where its main constraint is best served, and some can move depending on the software backend and how much delay the task can tolerate.
| Workload | Typical placement | Why | When it may move |
|---|---|---|---|
| Real-time rendering and viewport (TouchDesigner, Unreal Engine) | RTX | Frame budget, display outputs, GPU APIs | — |
| GPU rendering (Blender, Houdini and similar) | RTX | Discrete GPU throughput | — |
| CUDA-first image generation (for example ComfyUI pipelines) | RTX | CUDA ecosystem and GPU speed | When live rendering needs the GPU at the same time and the node’s backend supports the pipeline |
| CUDA-based training and fine-tuning | RTX | CUDA tooling | — |
| Low-latency vision (pose, depth or segmentation feeding the render) | RTX | Needs to stay close to the render loop | When the task tolerates delay and the node’s backend supports the model |
| Speech recognition | Depends | Backend and latency requirement | Often fine on the node when a short delay is acceptable |
| Large language models | Node | Memory capacity | Small models can stay on the RTX side if they fit and do not disturb the work |
| Long-context inference and retrieval | Node | Memory for context | — |
| Persistent agent backend | Node | Always available without occupying the creative GPU | — |
| Other memory-capacity-heavy workloads | Node | Unified memory capacity | — |
ComfyUI, speech recognition and vision models can move between machines depending on the software backend and the latency each task can tolerate.
For the general, non-creative version of this table, see What Runs Where.
Asynchronous Integration: Keep Inference Off the Render Loop
Do not make model inference a blocking dependency of the real-time render loop.
The pattern is the same in every tool:
- Send a request to the model server on the AI node.
- Keep rendering while the node works.
- Receive the response through a callback or event.
- Apply the result at a safe point, such as the next frame or a scene transition.
- Fall back gracefully if the response is late or missing.
Where the model server supports streaming, partial output can be shown as it arrives instead of waiting for the full response.
TouchDesigner
Derivative’s built-in operators cover this pattern. The Web Client DAT sends HTTP requests to the node’s API, with responses handled in its callbacks. The WebSocket DAT suits persistent two-way connections and streamed messages, and the SocketIO DAT applies when the server uses Socket.IO. For the real-time side of a TouchDesigner system, see the real-time section of the Creator guide.
Unreal Engine
Use the engine’s HTTP module: create an asynchronous HTTP request, bind a completion callback, and act on the result through an event. Do not block the game thread while waiting for AI inference. For how GPU and CPU loads divide in Unreal development, see the Unreal Engine section of the Creator guide.
Blender, Houdini and other DCC tools
The same principle applies: keep requests off the user-interface thread and bring results back as data — text, parameters or files. Heavy simulation and rendering workloads have their own sizing needs; see the Houdini workstation guide.
Latency: When Inference Speed Dominates
On a wired LAN, network latency is usually a small fraction of a model’s response time. The delay you notice comes mostly from the node: how long it takes to produce the first token, and how fast it generates the rest. For text, link bandwidth is not the deciding factor; see 2.5GbE vs 10GbE for Local AI for where bandwidth does matter.
Suits a two-PC setup
- Dialogue or characters that can tolerate a brief pause
- Scene, mood or parameter changes every few seconds
- Cue, caption or text generation
- Creator assistants and agents during production
- Batch preparation before a show or session
Does not suit it
- Decisions that must complete within one frame
- Tight synchronization to audio or video timing
- Anything the render cannot proceed without
Design tactics that help: pre-generate what you can, cache repeated responses, stream partial output, and define fallback states the visuals can hold while waiting.
Where ComfyUI and Image Generation Belong
By default, image generation stays on the RTX side. Most pipelines are CUDA-first and benefit directly from discrete GPU speed.
The catch is that generating images on the same GPU during live rendering brings contention back. Practical options:
- Schedule generation outside live sessions or between scenes.
- Accept the trade-off when the real-time load is light.
- Move some jobs to the node when its software backend supports the workflow and slower generation is acceptable.
Choosing the Node for a Creative Studio
Mac Studio
Quiet and compact, designed to remain quiet under sustained workloads, which suits studios and installation spaces. Offers very high unified-memory capacity and runs macOS with Apple’s MLX framework.
Ryzen AI Max systems
Supported Windows and Linux configurations are available; support remains platform- and configuration-dependent. Available in several form factors from different makers, so noise and ports depend on the specific system.
The AI node can run macOS, Windows or Linux depending on the node platform and software stack; it does not need to match the RTX workstation. For a full comparison, see Mac Studio vs Ryzen AI Max, and for current chips and memory configurations, its Current Platform Snapshot.
Pairing Tiers by Creative Role
Pairings are described by role and capacity tier rather than by product generation. Cost generally increases with memory tier, platform and form factor.
| Pairing | Host role | Node tier | Typical creative workflow |
|---|---|---|---|
| Real-Time Host + Mainstream Node | RTX workstation sized for scene complexity and output count | Mainstream (around 64GB) | Interactive installations or real-time development with a mid-size language model for dialogue, cues or parameters |
| CUDA-Heavy Host + High-Capacity Node | RTX workstation prioritizing GPU throughput for rendering, image generation and CUDA tools | High-capacity (around 128GB) | Daily use of large language models or long context alongside heavy GPU production |
| Production Workstation + Workstation-Class Node | Production-grade RTX workstation with expansion, I/O and sustained cooling | Workstation class (beyond 128GB) | Studio pipelines, several clients, very large models or a shared agent backend |
For pairings expressed by VRAM class, see Tier Pairings.
Decision Table
| If you… | Then |
|---|---|
| Run small models alongside light scenes and see no contention | One RTX PC is enough |
| See frame pacing or VRAM problems when AI runs during real-time work | Move large-model inference to an AI node |
| Need large language models or long context available all day while creating | RTX workstation plus a High-capacity node |
| Need image generation and live rendering at the same time | Keep generation on the RTX side, schedule it, or add GPU capacity; the node helps only if its backend fits the pipeline |
| Need responses within a single frame | Not a task for a networked language model; use on-GPU real-time techniques |
Frequently Asked Questions
Can I run a local LLM and TouchDesigner or Unreal Engine on the same RTX GPU?
Yes, if the model is small and both fit comfortably in VRAM. They still share VRAM, compute and memory bandwidth, so heavier scenes or larger models can affect frame pacing and may force models to be unloaded. Moving large-model inference to a separate node reduces that contention.
Should ComfyUI run on the RTX PC or the AI node?
Usually on the RTX PC, because most image-generation pipelines are CUDA-first and benefit from discrete GPU speed. Placement can change: if live rendering needs the GPU at the same time, or if the node’s software backend supports the workflow and slower generation is acceptable, the node can take some jobs.
Does a two-PC setup make AI responses fast enough for real-time visuals?
It protects the render loop, but it does not make the model faster. Response time is set mainly by the node’s inference speed. Two-PC setups suit interactions that tolerate a short delay, such as dialogue, cues and parameter changes, not decisions that must complete within a single frame.
How should creative software talk to the AI node?
Through an asynchronous network API: send a request, keep rendering, and apply the result when it arrives. In TouchDesigner this typically means the Web Client DAT or WebSocket DAT; in Unreal Engine, the engine’s HTTP module with a completion callback. Do not wait for inference on the render or game thread.
Which operating system should the AI node run?
macOS, Windows or Linux, depending on the node platform and software stack. Mac Studio runs macOS. For Ryzen AI Max systems, supported Windows and Linux configurations are available; support remains platform- and configuration-dependent. Choose the OS your model server and tools support best. The RTX workstation does not need to match.
Official Technical References
Technical references only. These links support the specifications and concepts on this page and are not purchase links.
- TouchDesigner: Web Client DAT
- TouchDesigner: WebSocket DAT
- TouchDesigner: SocketIO DAT
- Unreal Engine: HTTP module API reference, including asynchronous request completion
- ComfyUI: Official documentation
- AMD: Use ROCm on Radeon and Ryzen (Windows and Linux)
- Apple: MLX framework documentation
Next Steps
Plan the hardware pairing, or connect the two machines: