Skip to main content

[ 00 / 10 ]

[ PAITON | AMD GPU OPTIMIZATION | REAL WORKLOADS ]

MORE FROM AMD GPUS. PROVEN ON REAL MODELS.

Get more performance
from AMD GPUs.

Paiton finds what slows your model down, replaces the generic runtime path,
and accelerates language, image and video workloads without retraining.

[ 00 / 10 ]

[ LATEST BREAKTHROUGHS ]

MEASURED ON REAL GENERATION AND INFERENCE PATHS.

Benchmarks that make AMD
worth a second look

The gap is rarely the model. It is usually the runtime.
Paiton tunes the exact path your workload uses, then proves the gain with benchmarks.

Video diffusion: MI355X vs B200

17.6%

Wan2.2-T2V-A14B ran faster on AMD MI355X

5.501sAMD MI355X
6.672sNVIDIA B200

Paiton-Diffusers removed runtime waste from the generation path and let the hardware show up.

Read Wan2.2 Benchmark
Why it matters

You do not need a new model to improve unit economics. You need the runtime, kernels, and deployment path tuned around the workload you already run.

[ 00 / 10 ]

[ GPU ECONOMICS MODEL ]

TRANSLATE THROUGHPUT INTO REUSABLE CAPACITY.

EDITABLE PLANNING SCENARIO

Turn runtime gains
into a business case.

Faster inference is more than a benchmark number. It can reduce the GPU time required for the same workload,

or increase the output of hardware you already own.

Start with this illustrative scenario, then replace the assumptions with the uplift measured during your Paiton benchmark.

YOUR GPU SCENARIOAll four assumptions are editable
SAME WORKLOAD13.0% fewer GPU-hours per unit of output
16,000 current GPU-hours / month
Current runtime baselinePaiton-adjusted estimate: 13,913 GPU-hours

CAPACITY RELEASED

2,087GPU-hours / month

Equivalent to roughly 4.2 GPUs at this utilization level.

Potential compute-capacity value€6,261 / month€75,130 / year

SAME COMPUTE BUDGET

+15.0%potential output

Use the released runtime for more tokens, generations, jobs, or customer demand without expanding the GPU budget.

A published Radeon AI PRO R9700 test measured 39.77 versus 32.86 output tokens per second on the same Qwen3.8 workload. Gains remain workload-specific and that result is not applied automatically here.

View the R9700 benchmark
Benchmark my workload

Illustrative planning model. Actual gains depend on the model, batch size, sequence length, precision, runtime, hardware, scaling behaviour, and deployment. The calculation assumes throughput improvement translates proportionally into lower GPU runtime for the same workload. Released capacity may become a cash saving in usage-based infrastructure or reusable capacity on owned hardware. Figures exclude Paiton fees and do not guarantee savings or performance.

[ 00 / 10 ]

[ WHAT PAITON CHANGES ]

LANGUAGE, IMAGES, VIDEO WITH SOUND, AND CUSTOM STACKS.

Paiton workload optimization visual

Turn AMD hardware
into production throughput.

Whether you run LLMs, image and video diffusion, MoE, or custom architectures,
Paiton starts where money is lost: latency, throughput, memory pressure, and cost per useful output.

Performance icon

More tokens, less waiting

Serve more work from one GPU. The latest Qwen3.8 run reached 400.7 aggregate output tok/s across eight concurrent requests on R9700. Throughput across users is distinct from the speed of a single response.

Weights unchanged icon

No retraining required

Keep the model and improve the path around it with AMD-aware operators and custom .so files.

Runtime icon

Runtime work that pays back

ComfyUI workflows for FLUX.2 klein images and MiniMax H3 video with sound on R9700, Wan2.2 video on Instinct, and vLLM for language inference.

Language model icon

Language Models

Llama, Qwen, Deepseek, Gemma, Mistral, and CodeLlama with vLLM support, FP8 precision, and custom kernels when the gain is worth it.

Diffusion and video icon

Diffusion & Video

MiniMax H3 with native audio, Wan2.2, FLUX.2 klein, Stable Diffusion, SDXL and ControlNet, with workload-specific runtime and memory tuning.

Advanced model icon

Advanced Models

MoE, MLA, multimodal systems, proprietary architectures, expert kernels, novel attention, and custom inference stacks.

[ 00 / 10 ]

[ ENGAGEMENT FLOW ]

MEASURE. TUNE. PROVE.

A practical optimization sprint

Bring the workload, the target GPU, and the performance goal.
We turn that into a focused path from measurement to production-ready gain.

01Measure

Run the real workload and identify where latency, memory pressure, or cost per useful output is leaking value.

02Tune

Build the AMD-specific path with custom .so files, FP8 precision, kernel fusion, and the same model weights.

03Prove

Deploy the optimized path on AMD Instinct or Radeon AI PRO hardware and compare the result against the baseline.

[ 00 / 10 ]

[ DEPLOYMENT FIT ]

ROCM, KERNELS, MULTI-GPU, AND RUNTIME WORK.

Serious AMD tuning,
for production inference.

Runtime work

vLLMPaiton-DiffusersComfyUI (FLUX.2 klein and MiniMax H3 on R9700)ROCmHIPIndependent kernels

Performance levers

FP8 precisionTensor parallelismData parallelismKernel fusionCustom AMD operatorsMulti-GPU scaling

What gets tuned

FP8 can reduce memory pressure while preserving accuracy. Tensor parallelism spreads large models like Llama-3.1-405B across multiple AMD GPUs. Custom kernels target the bottlenecks generic runtimes leave behind.

CDNA 4.0

MI355X

288GB HBM3E

CDNA 3.0

MI325X / MI300X / MI300A

256GB HBM3E, 192GB HBM3, 128GB HBM3 APU

CDNA 2.0

MI250X / MI250 / MI210

Datacenter AMD GPU support

CDNA 1.0

MI100

First compute-optimized generation

RDNA 4.0

Radeon AI PRO R9700

Qualified Qwen3.8, Ornith 1.5, FLUX.2 klein and MiniMax H3 profiles on 32GB GDDR6

RDNA 3.0

RX 7900 XTX / RX 7900 XT

Supported for selected inference workloads

RDNA 2.0

RX 6800 XT / RX 6900 XT

Supported with workload-specific qualification

Paiton supports both AMD Instinct datacenter accelerators and Radeon GPUs. Every model, runtime, precision, and hardware combination is qualified against a defined deployment scope before delivery.

[ 00 / 10 ]

[ COMMUNITY ]

PRACTICAL LOCAL AI, SHARED WITH THE COMMUNITY.

Faster local AI.Shared with the community.

Better inference should reach beyond our customer projects. Alongside our open-source vLLM plugin, our free RDNA community packages help developers run language models, generate images and create video with sound on their own hardware.

Explore the code, run the supported models locally, and share what you learn. It is our way of helping more people get more from the GPUs they already own.

Plan your on-prem AI
Community packagesvLLM plugin: Apache-2.0

paiton-vllm-plugin

vLLM integration, 65K throughput and 200K chat presets, ComfyUI workflows and model-specific setup guides.

Explore the packages on GitHub

Qwen3.8 includes structured tool calls and a public experimental chat profile with prefix caching and a 220K context option. Long-context chat uses one active request and has tight VRAM headroom; release presets default to caching off. Separate packages cover FLUX.2 klein images and MiniMax H3 video with sound. Check each guide for limits and licenses. The Paiton compiler remains private.

[ 00 / 10 ]

[ EVIDENCE ]

PUBLISHED BENCHMARKS AND CASE STUDIES.

Read the benchmarks
behind the claims

400.7 aggregate tok/s

Qwen3.8 on ROCm 10: throughput and long conversations

New 65K-preset results, public 200K/220K chat instructions and experimental prefix caching. The article also preserves the original 57% gain against matched Radiance + DFlash2 in the separate 188-request, 8K-context comparison, plus the demo recording.

Benchmark + demo

16.7% less waiting

MiniMax H3 video and audio on one Radeon

A 15-second clip with native stereo audio, saved in 5m 33s on average versus 6m 39s for matched stock. Turbo8 on one 32 GB R9700, with warm, complete-request timings.

Watch the clips

16.2% less generation time

FLUX.2 klein on Radeon AI PRO R9700

1.054 seconds per image with a warm pipeline, 33.4% lower peak Torch allocation, and a free local ComfyUI workflow. Tested at 1024 × 1024, four steps and batch one.

Read Benchmark

+27% output

Ornith 1.5 on one 32GB Radeon GPU

Paiton with DFlash delivers 44.63 output tokens per second on R9700, 27% ahead of tuned stock vLLM in the tested single-user workload.

Read Benchmark

+21% output

Qwen3.8 on Radeon AI PRO R9700

Higher interactive, coding, and long-context throughput from the same 32GB Radeon GPU.

Read Benchmark

405B model

Llama-3.1-405B acceleration

Faster startup and inference for a frontier-scale open model.

Read Case Study

$/1M tokens

MoE kernels beat H200/B200

MI300X plus Paiton runtime on long-context MoE economics.

Read Benchmark

[ 00 / 10 ]

[ NEXT STEP ]

START WITH THE MODEL AND TARGET GPU.
Paiton CTA background accent

Show us the workload.
We will show you the gain.

Share the model, target GPU, and production goal.
Paiton turns that into a benchmark-led optimization plan.