[{"data":1,"prerenderedAt":1467},["ShallowReactive",2],{"blog-post-en-\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700":3,"blog-posts-sidebar-en":992},{"id":4,"title":5,"body":6,"categories":973,"date":979,"description":980,"extension":981,"heading":982,"image":983,"meta":984,"navigation":985,"originalUrl":986,"path":986,"seo":987,"slug":988,"stem":989,"updated":990,"__hash__":991},"blog\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700.md","Local FLUX.2 klein on Radeon AI PRO R9700: Faster Image Generation with Less VRAM",{"type":7,"value":8,"toc":961},"minimark",[9,13,21,44,55,61,67,72,75,140,166,173,180,187,196,200,203,209,258,269,280,284,287,290,296,350,356,359,366,370,385,391,394,397,403,465,468,479,483,486,525,532,543,546,550,553,575,588,594,601,605,612,615,741,744,751,755,776,779,797,801,804,844,864,867,920,932,936,948,957],[10,11,12],"p",{},"A product concept. A woodland photograph. A rainy street in watercolor. The useful part of local image generation is being able to change the prompt and try again, without sending the request to an external service.",[10,14,15,16,20],{},"Our free Paiton profile brings that workflow to ",[17,18,19],"strong",{},"FLUX.2 klein 4B on one AMD Radeon AI PRO R9700",", with a ready-to-run ComfyUI setup.",[10,22,23,24,27,28,31,32,35,36,39,40,43],{},"At ",[17,25,26],{},"1024 × 1024, four steps and batch size one",", complete warm generation averages ",[17,29,30],{},"1.054 seconds per image",", against ",[17,33,34],{},"1.258 seconds"," for the strongest stock configuration we qualified. That is ",[17,37,38],{},"16.2% lower generation latency"," and ",[17,41,42],{},"19.4% more projected images per hour",".",[10,45,46,47,50,51,54],{},"Memory use falls too: peak Torch allocation drops from ",[17,48,49],{},"19.3 to 12.9 GiB",", a reported ",[17,52,53],{},"33.4% reduction",". The text encoder, transformer and VAE stay on the GPU, with no CPU offload.",[10,56,57,60],{},[17,58,59],{},"The headline timing is warm prompt-to-image generation, not startup or browser-click latency."," It includes text encoding, denoising, VAE decoding and conversion to a PIL image; it excludes PNG encoding, file writing and the interface.",[10,62,63],{},[64,65,66],"em",{},"The fox pictured above is an actual retained benchmark output from the Paiton FLUX.2 klein profile, seed 42.",[68,69,71],"h2",{"id":70},"start-creating-locally","Start creating locally",[10,73,74],{},"On a Linux R9700 workstation with Docker, the Compose plugin and a working AMD GPU driver, launch the included workflow with:",[76,77,82],"pre",{"className":78,"code":79,"language":80,"meta":81,"style":81},"language-bash shiki shiki-themes github-light github-dark","git clone --depth 1 \\\n  --branch paiton-flux2-klein-gfx1201-v1.0.1 \\\n  https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\ncd paiton-vllm-plugin\n.\u002Fmodels\u002FFLUX.2-klein\u002Flaunch.sh\n","bash","",[83,84,85,108,119,125,134],"code",{"__ignoreMap":81},[86,87,90,94,98,102,105],"span",{"class":88,"line":89},"line",1,[86,91,93],{"class":92},"sScJk","git",[86,95,97],{"class":96},"sZZnC"," clone",[86,99,101],{"class":100},"sj4cs"," --depth",[86,103,104],{"class":100}," 1",[86,106,107],{"class":100}," \\\n",[86,109,111,114,117],{"class":88,"line":110},2,[86,112,113],{"class":100},"  --branch",[86,115,116],{"class":96}," paiton-flux2-klein-gfx1201-v1.0.1",[86,118,107],{"class":100},[86,120,122],{"class":88,"line":121},3,[86,123,124],{"class":96},"  https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\n",[86,126,128,131],{"class":88,"line":127},4,[86,129,130],{"class":100},"cd",[86,132,133],{"class":96}," paiton-vllm-plugin\n",[86,135,137],{"class":88,"line":136},5,[86,138,139],{"class":92},".\u002Fmodels\u002FFLUX.2-klein\u002Flaunch.sh\n",[10,141,142,143,150,151,154,155,158,159,162,163,43],{},"Open ",[144,145,149],"a",{"href":146,"rel":147},"http:\u002F\u002F127.0.0.1:8188\u002F?paiton=1",[148],"nofollow","ComfyUI at localhost:8188",", edit the prompt and click ",[17,152,153],{},"Run",". The connected workflow includes an engine selector for ",[17,156,157],{},"Paiton"," or ",[17,160,161],{},"Stock (Diffusers)",", a seed control, an image preview and a Save Image node. Images are saved in ",[83,164,165],{},"paiton-images\u002F",[10,167,168],{},[169,170],"img",{"alt":171,"src":172},"The included ComfyUI workflow, showing the Paiton engine selector, prompt and image preview","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fcomfyui-workflow.webp",[10,174,175,176,179],{},"The first image loads and compiles the selected engine and can take several minutes. Later prompts reuse it. Switching engines unloads the previous model and incurs setup again; they are not held in GPU memory together. To compare outputs, keep the prompt and seed unchanged and select ",[17,177,178],{},"fixed"," in the seed control.",[10,181,182,183,186],{},"The tested host has 16 GB of RAM. We recommend ",[17,184,185],{},"24 GB of system RAM and 60 GB of free disk space"," for compilation headroom, containers, model files, caches and outputs. Downloads and caches persist between starts. Once setup is complete, generation stays local; no paid inference service or cloud GPU is required.",[10,188,189,190,195],{},"The ",[144,191,194],{"href":192,"rel":193},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Ftree\u002Fmain\u002Fmodels\u002FFLUX.2-klein",[148],"model guide"," also covers a lightweight prompt interface, terminal generation and adding the node to an existing ComfyUI installation.",[68,197,199],{"id":198},"what-the-speed-result-includes","What the speed result includes",[10,201,202],{},"Both engines use the same model checkpoint and generation settings. The comparison measures the whole warm generation loop rather than extrapolating from a faster individual kernel.",[10,204,205],{},[169,206],{"alt":207,"src":208},"Warm generation takes 1.054 seconds with Paiton versus 1.258 seconds with qualified stock, with projected hourly output of 3,416 versus 2,862 images","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fgeneration-performance.webp",[210,211,212,228],"table",{},[213,214,215],"thead",{},[216,217,218,222,226],"tr",{},[219,220,221],"th",{},"Generation metric",[219,223,225],{"align":224},"right","Qualified stock + weight cache",[219,227,157],{"align":224},[229,230,231,245],"tbody",{},[216,232,233,237,240],{},[234,235,236],"td",{},"Mean warm seconds per image",[234,238,239],{"align":224},"1.258",[234,241,242],{"align":224},[17,243,244],{},"1.054",[216,246,247,250,253],{},[234,248,249],{},"Projected images per hour",[234,251,252],{"align":224},"2,862",[234,254,255],{"align":224},[17,256,257],{},"3,416",[10,259,260,261,264,265,268],{},"Projected output is ",[83,262,263],{},"3600 \u002F mean generation seconds",". It is ",[17,266,267],{},"not an hour-long throughput measurement",", and it excludes PNG writing and time spent editing prompts. The 19.4% output increase and 16.2% latency reduction express the same improvement in different ways.",[10,270,271,272,275,276,279],{},"The interface has its own costs. In a separate ComfyUI validation, two warm workflow executions took ",[17,273,274],{},"1.408 and 1.404 seconds",", averaging ",[17,277,278],{},"1.406 seconds"," including local engine transport, image handling and saving. Those server-side timings exclude browser display and are not the headline prompt-to-PIL benchmark.",[68,281,283],{"id":282},"lower-memory-use-without-cpu-offload","Lower memory use, without CPU offload",[10,285,286],{},"A model's download size is not its runtime memory requirement. The full pipeline also needs its text encoder, decoder, activations, temporary buffers and compilation allocations.",[10,288,289],{},"Paiton reduces all three reported memory measures in this comparison:",[10,291,292],{},[169,293],{"alt":294,"src":295},"Three separate memory measurements for stock and Paiton: peak Torch allocation 19.3 versus 12.9 GiB, peak reservation 22.1 versus 14.1 GiB, and maximum sampled driver VRAM 23.0 versus 14.6 GiB","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fpipeline-memory.webp",[210,297,298,309],{},[213,299,300],{},[216,301,302,305,307],{},[219,303,304],{},"Memory measurement",[219,306,225],{"align":224},[219,308,157],{"align":224},[229,310,311,324,337],{},[216,312,313,316,319],{},[234,314,315],{},"Peak Torch allocation",[234,317,318],{"align":224},"19.3 GiB",[234,320,321],{"align":224},[17,322,323],{},"12.9 GiB",[216,325,326,329,332],{},[234,327,328],{},"Peak Torch reservation",[234,330,331],{"align":224},"22.1 GiB",[234,333,334],{"align":224},[17,335,336],{},"14.1 GiB",[216,338,339,342,345],{},[234,340,341],{},"Maximum sampled driver VRAM",[234,343,344],{"align":224},"23.0 GiB",[234,346,347],{"align":224},[17,348,349],{},"14.6 GiB",[10,351,352,355],{},[17,353,354],{},"12.9 GiB is not total VRAM consumption."," Allocation records tensor memory; reservation records memory held by the Torch allocator; driver samples include memory outside that allocator. These views overlap and must not be added together.",[10,357,358],{},"Allocation and reservation peaks include compilation, warmup and generation. Driver sampling also includes loading. Values are rounded and use binary GiB; the card is marketed as 32 GB.",[10,360,361,362,365],{},"The practical benefit is more memory headroom on the tested R9700 while retaining GPU-resident generation. ",[17,363,364],{},"These measurements do not establish support for a 16 GB GPU or any other card."," This release qualifies the R9700 only.",[68,367,369],{"id":368},"same-settings-inspectable-images-not-identical-pixels","Same settings, inspectable images, not identical pixels",[10,371,372,373,376,377,380,381,384],{},"The retained examples cover wildlife photography at seed ",[17,374,375],{},"42",", product photography at ",[17,378,379],{},"31415",", and a watercolor bookshop at ",[17,382,383],{},"2026",". All final latents were finite, and each fixed-seed result repeated exactly within its benchmark process.",[10,386,387],{},[169,388],{"alt":389,"src":390},"The retained stock and Paiton product images, with a teal mug, lemon and PAITON card; both use seed 31415","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fproduct-comparison.webp",[10,392,393],{},"Both product images retain the teal cup, yellow lemon, window light and readable “PAITON” card. The fox pair preserves the overall pose and woodland lighting. Fine textures, reflections and shadows differ.",[10,395,396],{},"The bookshop pair shows a larger difference in scene details, including masonry, shelving and lettering. Both retain the watercolor style, wet street, red raincoat and bicycle. The person stands beside the bicycle rather than visibly riding it.",[10,398,399],{},[169,400],{"alt":401,"src":402},"The retained stock and Paiton watercolor bookshop images at seed 2026, showing similar composition with differences in scene details","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fbookshop-comparison.webp",[210,404,405,421],{},[213,406,407],{},[216,408,409,412,415,418],{},[219,410,411],{},"Prompt",[219,413,414],{"align":224},"RGB SSIM",[219,416,417],{"align":224},"RGB PSNR",[219,419,420],{"align":224},"Final latent relative RMSE",[229,422,423,437,451],{},[216,424,425,428,431,434],{},[234,426,427],{},"Fox",[234,429,430],{"align":224},"0.975",[234,432,433],{"align":224},"32.45 dB",[234,435,436],{"align":224},"0.146",[216,438,439,442,445,448],{},[234,440,441],{},"Product",[234,443,444],{"align":224},"0.974",[234,446,447],{"align":224},"30.10 dB",[234,449,450],{"align":224},"0.242",[216,452,453,456,459,462],{},[234,454,455],{},"Bookshop",[234,457,458],{"align":224},"0.859",[234,460,461],{"align":224},"21.38 dB",[234,463,464],{"align":224},"0.289",[10,466,467],{},"These metrics measure similarity to the stock output, not aesthetic quality or complete prompt adherence. Small numerical differences can propagate through denoising, so operator checks are accompanied by final-image inspection.",[10,469,470,471,474,475,478],{},"Repeatability also has a boundary: a separately compiled, empty-cache Paiton run produced a fox image with SSIM ",[17,472,473],{},"0.948"," against the populated-cache Paiton image. The cause was not isolated to one operation. This three-prompt check does ",[17,476,477],{},"not"," establish unchanged quality or identical pixels across all prompts, styles or independently compiled processes.",[68,480,482],{"id":481},"what-to-expect-at-startup","What to expect at startup",[10,484,485],{},"Once the engine is loaded, repeated generation is the fast path. A new process still pays for loading and graph initialization, even with persistent caches.",[210,487,488,501],{},[213,489,490],{},[216,491,492,495,498],{},[219,493,494],{},"Paiton startup condition",[219,496,497],{"align":224},"Process loading",[219,499,500],{"align":224},"First generation",[229,502,503,514],{},[216,504,505,508,511],{},[234,506,507],{},"Compilation caches populated",[234,509,510],{"align":224},"44.7 s",[234,512,513],{"align":224},"19.4 s",[216,515,516,519,522],{},[234,517,518],{},"Compilation caches empty; weights already prepared",[234,520,521],{"align":224},"48.2 s",[234,523,524],{"align":224},"104.1 s",[10,526,527,528,531],{},"The first-generation timings include graph setup and any remaining compilation. The next two warm generations in the empty-cache test took ",[17,529,530],{},"1.050 and 1.052 seconds",". These startup examples are separate from the six-image headline comparison.",[10,533,534,535,538,539,542],{},"Initial provisioning adds the download and model-preparation stages. On our system, the source download took ",[17,536,537],{},"110.7 seconds"," and conversion took ",[17,540,541],{},"88.7 seconds",". Network speed and cache history will change those numbers.",[10,544,545],{},"For repeated use, keep the selected engine loaded and iterate on prompts rather than switching backends after every image.",[68,547,549],{"id":548},"more-images-per-euro","More images per euro",[10,551,552],{},"At the same cost per productive hour, lower generation time gives the workstation more output capacity.",[10,554,555,556,559,560,563,564,567,568,571,572,43],{},"For an illustrative ownership model, assume ",[17,557,558],{},"€2,000"," for the workstation, ",[17,561,562],{},"4,000 productive generation hours",", electricity at ",[17,565,566],{},"€0.30\u002FkWh",", and equal assumed wall power of ",[17,569,570],{},"350 W"," for both engines. Hardware contributes €0.50 per hour and electricity adds €0.105, giving ",[17,573,574],{},"€0.605 per productive hour",[10,576,577,578,39,581,584,585,43],{},"At the reported generation rates, that produces approximately ",[17,579,580],{},"4,730 images per euro for stock",[17,582,583],{},"5,647 for Paiton",", ",[17,586,587],{},"19.4% more modeled output per euro",[10,589,590],{},[169,591],{"alt":592,"src":593},"Modeled generation capacity at the same assumed hourly cost: 4,730 images per euro for stock and 5,647 for Paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Fmodeled-cost.webp",[10,595,596,597,600],{},"This is a capacity and cost model, ",[17,598,599],{},"not measured energy savings",". It assumes sustained productive use and excludes idle time, taxes, labor, financing, cooling outside the workstation and PNG writing. It also does not measure how many generated images a user will choose to keep. Lower utilization raises capital cost per useful image.",[68,602,604],{"id":603},"how-the-comparison-was-run","How the comparison was run",[10,606,607,608,611],{},"The baseline was the strongest stock path retained during our qualification, not an untouched default installation. It used ",[17,609,610],{},"Diffusers 0.40.0 and SDNQ 0.2.6",", prepared immutable weights, compiled pipeline components and graph capture. Other tuning options were tested; not every option improved the full pipeline.",[10,613,614],{},"Both engines used the same pinned checkpoint, prompts, fixed seeds, equivalent runtime precision and full GPU residency. Paiton contributes a qualified compiled execution path for this model and Radeon profile. The scheduler, step count and guidance remain unchanged; the result is not obtained by generating a smaller image or running fewer steps.",[210,616,617,627],{},[213,618,619],{},[216,620,621,624],{},[219,622,623],{},"Tested setting",[219,625,626],{},"Value",[229,628,629,637,645,653,661,669,677,685,693,701,711,721,731],{},[216,630,631,634],{},[234,632,633],{},"GPU",[234,635,636],{},"AMD Radeon AI PRO R9700, 32 GB",[216,638,639,642],{},[234,640,641],{},"Model profile",[234,643,644],{},"FLUX.2 klein 4B, text-to-image",[216,646,647,650],{},[234,648,649],{},"Resolution \u002F steps \u002F batch",[234,651,652],{},"1024 × 1024 \u002F 4 \u002F 1",[216,654,655,658],{},[234,656,657],{},"Guidance \u002F text sequence",[234,659,660],{},"1.0 \u002F 512 tokens",[216,662,663,666],{},[234,664,665],{},"Measurement date",[234,667,668],{},"September 7, 2026",[216,670,671,674],{},[234,672,673],{},"Repetitions",[234,675,676],{},"Three fixed prompts; two warmups and two measured runs per prompt",[216,678,679,682],{},[234,680,681],{},"Measured sample",[234,683,684],{},"Six images per backend",[216,686,687,690],{},[234,688,689],{},"Timing boundary",[234,691,692],{},"GPU-synchronized wall clock, prompt to PIL image",[216,694,695,698],{},[234,696,697],{},"GPU policy",[234,699,700],{},"AUTO performance level, COMPUTE profile",[216,702,703,706],{},[234,704,705],{},"Common PyTorch build",[234,707,708],{},[83,709,710],{},"2.12.0+rocm7.14.0",[216,712,713,716],{},[234,714,715],{},"HIP",[234,717,718],{},[83,719,720],{},"7.14.60850",[216,722,723,726],{},[234,724,725],{},"Triton",[234,727,728],{},[83,729,730],{},"3.7.1+git0263a6a6.rocm7.14.0",[216,732,733,736],{},[234,734,735],{},"Transformers",[234,737,738],{},[83,739,740],{},"5.15.1",[10,742,743],{},"No clock, power, voltage or fan limits were changed. Automatic clocks and temperatures varied. This small sample describes the tested workstation and settings; it is not a universal speed guarantee.",[10,745,746,747,750],{},"The supported release is deliberately specific: ",[17,748,749],{},"Linux, R9700, text-to-image, 1024-square, four steps, batch one",". Image editing, adapters, other resolutions and other GPUs are outside its qualified scope.",[68,752,754],{"id":753},"model-provenance-and-release-notes","Model provenance and release notes",[10,756,757,758,763,764,767,768,771,772,775],{},"The package downloads ",[144,759,762],{"href":760,"rel":761},"https:\u002F\u002Fhuggingface.co\u002FDisty0\u002FFLUX.2-klein-4B-SDNQ-4bit-dynamic",[148],"Disty0's SDNQ checkpoint",", pinned to revision ",[83,765,766],{},"45e9cc76cb70f84473ce5c6c2e2282d0ef3c6ecd",". The download is about ",[17,769,770],{},"5.46 GB","; a separate preparation stage produces about ",[17,773,774],{},"12 GB"," of tensor files. Checkpoint size should not be confused with runtime precision or GPU memory usage.",[10,777,778],{},"The release notes identify the original FLUX.2 klein 4B weights as Apache 2.0. They also flag a contradictory non-commercial link in the community checkpoint's metadata and an unpinned pre-quantization source revision. Weights are downloaded separately; those provenance limitations should be reviewed before commercial redistribution.",[10,780,781,782,584,787,39,792,796],{},"The SDNQ conversion\u002Fstock tools and ComfyUI retain their upstream source and notices. The Paiton inference path does not import SDNQ. See the ",[144,783,786],{"href":784,"rel":785},"https:\u002F\u002Fhuggingface.co\u002Fblack-forest-labs\u002FFLUX.2-klein-4B",[148],"official model card",[144,788,791],{"href":789,"rel":790},"https:\u002F\u002Fgithub.com\u002FDisty0\u002Fsdnq",[148],"SDNQ source",[144,793,795],{"href":192,"rel":794},[148],"release guide"," for the associated documentation.",[68,798,800],{"id":799},"reproduce-the-workflow","Reproduce the workflow",[10,802,803],{},"From the cloned repository, enter the model directory and generate an image:",[76,805,807],{"className":78,"code":806,"language":80,"meta":81,"style":81},"cd models\u002FFLUX.2-klein\n.\u002Frun.sh generate \\\n  --prompt 'A teal ceramic coffee cup beside a lemon, soft window light, product photograph' \\\n  --seed 42\n",[83,808,809,816,826,836],{"__ignoreMap":81},[86,810,811,813],{"class":88,"line":89},[86,812,130],{"class":100},[86,814,815],{"class":96}," models\u002FFLUX.2-klein\n",[86,817,818,821,824],{"class":88,"line":110},[86,819,820],{"class":92},".\u002Frun.sh",[86,822,823],{"class":96}," generate",[86,825,107],{"class":100},[86,827,828,831,834],{"class":88,"line":121},[86,829,830],{"class":100},"  --prompt",[86,832,833],{"class":96}," 'A teal ceramic coffee cup beside a lemon, soft window light, product photograph'",[86,835,107],{"class":100},[86,837,838,841],{"class":88,"line":127},[86,839,840],{"class":100},"  --seed",[86,842,843],{"class":100}," 42\n",[10,845,846,847,850,851,854,855,858,859,43],{},"Use ",[83,848,849],{},"--count 4"," to try consecutive seeds in one loaded process. The default output is ",[83,852,853],{},"outputs\u002Fimage.png",". For the lightweight interface, run ",[83,856,857],{},".\u002Flaunch.sh --ui simple"," and open ",[144,860,863],{"href":861,"rel":862},"http:\u002F\u002F127.0.0.1:7860",[148],"localhost:7860",[10,865,866],{},"Stop other generation services before comparing both engines:",[76,868,870],{"className":78,"code":869,"language":80,"meta":81,"style":81},".\u002Flaunch.sh --stop\n.\u002Frun.sh benchmark --backend stock --suite --output \u002Foutputs\u002Fstock\n.\u002Frun.sh benchmark --backend paiton --suite --output \u002Foutputs\u002Fpaiton\n",[83,871,872,880,902],{"__ignoreMap":81},[86,873,874,877],{"class":88,"line":89},[86,875,876],{"class":92},".\u002Flaunch.sh",[86,878,879],{"class":100}," --stop\n",[86,881,882,884,887,890,893,896,899],{"class":88,"line":110},[86,883,820],{"class":92},[86,885,886],{"class":96}," benchmark",[86,888,889],{"class":100}," --backend",[86,891,892],{"class":96}," stock",[86,894,895],{"class":100}," --suite",[86,897,898],{"class":100}," --output",[86,900,901],{"class":96}," \u002Foutputs\u002Fstock\n",[86,903,904,906,908,910,913,915,917],{"class":88,"line":121},[86,905,820],{"class":92},[86,907,886],{"class":96},[86,909,889],{"class":100},[86,911,912],{"class":96}," paiton",[86,914,895],{"class":100},[86,916,898],{"class":100},[86,918,919],{"class":96}," \u002Foutputs\u002Fpaiton\n",[10,921,189,922,927,928,931],{},[144,923,926],{"href":924,"rel":925},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FFLUX.2-klein-4B-Paiton-RDNA4",[148],"compiled artifacts on Hugging Face"," are already included in the containers. The release guide covers the build recipes, runtime bindings, workflow, launchers and notices. Use ",[83,929,930],{},".\u002Flaunch.sh --build"," to build the containers locally.",[68,933,935],{"id":934},"more-useful-work-from-amd-hardware","More useful work from AMD hardware",[10,937,938,939,39,943,947],{},"Our recent ",[144,940,942],{"href":941},"\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","Qwen3.8",[144,944,946],{"href":945},"\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700","Ornith 1.5"," releases focused on local language-model serving. This profile brings the same practical focus to a visual workflow: faster iteration, less memory pressure and a ComfyUI setup people can use on their own machine.",[10,949,950,953,954,43],{},[17,951,952],{},"Building image-generation services or running inference at scale on AMD CDNA? Talk to us about your workload."," Paiton's wider work targets throughput, memory efficiency and cost per useful output across AMD deployments. Learn more about ",[144,955,157],{"href":956},"\u002Fproducts\u002Fpaiton",[958,959,960],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":81,"searchDepth":110,"depth":110,"links":962},[963,964,965,966,967,968,969,970,971,972],{"id":70,"depth":110,"text":71},{"id":198,"depth":110,"text":199},{"id":282,"depth":110,"text":283},{"id":368,"depth":110,"text":369},{"id":481,"depth":110,"text":482},{"id":548,"depth":110,"text":549},{"id":603,"depth":110,"text":604},{"id":753,"depth":110,"text":754},{"id":799,"depth":110,"text":800},{"id":934,"depth":110,"text":935},[157,974,975,976,977,978],"AMD Radeon","Local AI","Image Generation","FLUX","ComfyUI","2026-09-07T09:00:00","Paiton generates 1024×1024 FLUX.2 klein images in 1.054 seconds on an R9700: 16.2% lower warm latency and 33.4% lower peak Torch allocation.","md","Faster local FLUX.2 klein on Radeon, with less memory","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Ffox-paiton.webp",{},true,"\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700",{"title":5,"description":980},"paiton-flux2-klein-radeon-ai-pro-r9700","blog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700",null,"IkZeqGXhXUn8R6da2Mes4z3NTsvf5J6egbxxPmNIT2s",[993,1005,1015,1026,1028,1042,1050,1064,1095,1107,1129,1147,1166,1184,1202,1218,1230,1246,1261,1273,1282,1290,1305,1317,1328,1339,1349,1362,1372,1385,1396,1406,1417,1426,1438,1449,1458],{"path":994,"title":995,"description":996,"date":997,"slug":998,"image":999,"originalUrl":1000,"categories":1001},"\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700","Qwen3.8: 400.7 tok\u002Fs on R9700 | Paiton","Qwen3.8 on one R9700: 400.7 aggregate tok\u002Fs with ROCm 10 and vLLM 0.29, plus public 200K\u002F220K chat profiles. Benchmarks, limits and launch commands.","2026-09-16T07:30:00Z","paiton-qwen38-mxfp4-dflash2-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002Fupdate-2026-09-19\u002Fhero.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700",[157,974,1002,942,1003,1004],"vLLM","Inference Optimization","DFlash2",{"path":1006,"title":1007,"description":1008,"date":1009,"slug":1010,"image":1011,"originalUrl":1012,"categories":1013},"\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700","Qwen3.8 GGUF in vLLM: Faster Responses on One Radeon","Run the original NEO CODER MAX GGUF in vLLM with Paiton on an R9700. Explore measured latency gains, image input and local deployment.","2026-09-14T07:30:00Z","paiton-qwen38-neo-gguf-vllm-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-neo-gguf\u002F00-hero-neo-gguf-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700",[157,974,975,1014,1002],"GGUF",{"path":1016,"title":1017,"description":1018,"date":1019,"slug":1020,"image":1021,"originalUrl":1022,"categories":1023},"\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700","MiniMax H3 on Radeon: 15-Second Video With Native Sound","Paiton generates a 15-second MiniMax H3 video with stereo audio on one Radeon AI PRO R9700 in 5m 33s, with 16.7% lower latency than matched stock.","2026-09-09T07:30:00Z","paiton-minimax-h3-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002F00-featured-minimax-h3-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700",[157,974,975,1024,1025,978],"Video Generation","MiniMax H3",{"path":986,"title":5,"description":980,"date":979,"slug":988,"image":983,"originalUrl":986,"categories":1027},[157,974,975,976,977,978],{"path":945,"title":1029,"description":1030,"date":1031,"slug":1032,"image":1033,"originalUrl":1034,"categories":1035},"Ornith 1.5 at 44.6 tok\u002Fs on One Radeon AI PRO R9700","Paiton serves Ornith 1.5 35B A3B at 44.63 output tok\u002Fs on one Radeon AI PRO R9700, 27% faster than tuned stock vLLM, with 21.3% lower modeled cost.","2026-09-05T09:00:00","paiton-ornith15-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-ornith15\u002F00-featured-ornith15-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700",[157,1036,974,1037,1038,1039,1003,1040,1002,1041],"Artificial Intelligence","AI Inference","GPU Performance","Inference Latency","Large Language Models","Cost Efficiency",{"path":941,"title":1043,"description":1044,"date":1045,"slug":1046,"image":1047,"originalUrl":1048,"categories":1049},"Paiton: 21% More Qwen3.8 Throughput on Radeon AI PRO R9700","Paiton serves AMD’s Qwen3.8 27B at 39.77 output tokens\u002Fs on one Radeon AI PRO R9700, delivering 21% more throughput and 17.4% lower modeled cost.","2026-09-04T09:00:00","paiton-qwen38-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F00-featured-paiton-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700",[157,1036,974,1037,1038,1039,1003,1040,1002,1041],{"path":1051,"title":1052,"description":1053,"date":1054,"slug":1055,"image":1056,"originalUrl":990,"categories":1057},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",[1058,1059,1060,1061,1062,1063],"All","AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":1065,"title":1066,"description":1067,"date":1068,"slug":1069,"image":1070,"originalUrl":1071,"categories":1072},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Wan2.2 Video Generation: Paiton on AMD MI355X","Compare Wan2.2-T2V-A14B video generation on AMD MI355X with Paiton and NVIDIA B200 using Diffusers, and explore our diffusion optimization approach.","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[1058,1036,157,1073,1074,1075,1076,1077,1078,1079,1080,1081,633,1082,1083,1084,1085,1086,1087,1088,157,1089,1090,1091,1092,1093,1094],"14B","AMD","B200","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","Hardware","Inference","Instinct","MI355x","NVidia","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":1096,"title":1097,"description":1098,"date":1099,"slug":1100,"image":1101,"originalUrl":1102,"categories":1103},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","ElioVP in De Tijd: Chip Optimization and Data Centers","Read about De Tijd's coverage of ElioVP, from its origins in chip optimization to its work on modular data centers and high-density cooling.","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[1058,1036,1104,1105,1074,1106,1104,1086],"Modular DC","Uncategorized","De Tijd",{"path":1108,"title":1109,"description":1110,"date":1111,"slug":1112,"image":1113,"originalUrl":1114,"categories":1115},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","AI Privacy: A Strategic Priority for Benelux Businesses","Explore generative AI privacy risks, trust, data retention and governance, and why Benelux businesses need a strategic approach to secure AI.","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[1058,1036,1116,1105,1117,1118,1119,1120,1121,1122,1123,1124,1080,1125,1081,1126,1127,1128],"Trending","AI Act","Anthropomorphism","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Microsoft Copilot","Privacy","Shadow AI",{"path":1130,"title":1131,"description":1132,"date":1133,"slug":1134,"image":1135,"originalUrl":1136,"categories":1137},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Why We Do Not Use itsme: Privacy and Data Sovereignty","Why ElioVP does not use itsme: our assessment of identity metadata, cloud dependence, data sovereignty and authentication risks.","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[1058,1138,1139,1140,1122,1141,1142,1143,1125,1144,1145,1146,1127],"AWS","Belgian Mobile ID","Cloud Act","Data Sovereignty","Digital Identity","eIDAS","itsme","Liberty Global","MyGov.be",{"path":1148,"title":1149,"description":1150,"date":1151,"slug":1152,"image":1153,"originalUrl":1154,"categories":1155},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","Lessons from building on-premise AI agents in 2025 cover workflow design, observability, model training, hallucinations and GPU memory limits.","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[1058,1036,1156,1116,1157,1158,1159,1160,1161,1162,1163,1164,1089,1165],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":1167,"title":1168,"description":1169,"date":1170,"slug":1171,"image":1172,"originalUrl":1173,"categories":1174},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","An analysis of AI neocloud investment risks, examining circular financing, infrastructure claims, contract terms and due diligence.","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[1058,1036,1116,1059,1175,1176,1177,1178,1179,1180,1181,1182,1183],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":1185,"title":1186,"description":1187,"date":1188,"slug":1189,"image":1190,"originalUrl":1191,"categories":1192},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","NVIDIA GB300 NVL72: A Four-Month Modular Data Center Plan","Explore a modular data center design for NVIDIA GB300 NVL72, covering redundant power, hybrid cooling and a four-month deployment plan.","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[1058,1104,1105,1193,1059,1194,1195,1196,1197,1198,1199,1200,1201],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":1203,"title":1204,"description":1205,"date":1206,"slug":1207,"image":1208,"originalUrl":1209,"categories":1210},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Why CUDA compatibility is not the same as AMD performance: explore ROCm, HIP, kernel tuning and the case for hardware-specific optimization.","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[1058,1036,157,1105,1211,1036,1212,1213,1214,1215,715,1216,157,1217],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","Kernel Tuning","ROCm",{"path":1219,"title":1220,"description":1221,"date":1222,"slug":1223,"image":1224,"originalUrl":1225,"categories":1226},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Learn how Paiton integrates with existing inference stacks, with AMD MI300X benchmark results and performance-per-dollar comparisons.","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[1058,1036,157,1037,1227,1211,1041,1228,1003,1216,157,1229,1002],"AMD Instinct","High Throughput","SGLang",{"path":1231,"title":1232,"description":1233,"date":1234,"slug":1235,"image":1236,"originalUrl":1237,"categories":1238},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Paiton MoE Benchmarks: MI300X vs H200 and B200","Compare Qwen3-30B-A3B MoE inference with Paiton on MI300X against H200 and B200, including throughput and cost per million tokens.","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[1058,1036,157,1239,1211,1240,1003,1241,1242,1243,1244,157,1245],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":1247,"title":1248,"description":1249,"date":1250,"slug":1251,"image":1252,"originalUrl":1253,"categories":1254},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Local Agentic AI: From Inbox to Action","Local-first AI agents turn email, documents and images into tickets, reports and actions, using models tailored to your data and systems.","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[1058,1036,1156,1105,1157,1255,1256,1257,1258,1162,1164,1089,1259,1260],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":1262,"title":1263,"description":1264,"date":1265,"slug":1266,"image":1267,"originalUrl":1268,"categories":1269},"\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","MI300X FP8 Benchmarks: GPU Partitioning with Paiton","Explore Llama 3.1 8B FP8 benchmarks on partitioned MI300X GPUs with Paiton, comparing throughput and latency against NVIDIA H200 and B200.","2025-07-31T13:32:57","mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F",[1058,1036,157,1105,1270,1074,1075,1271,1272,1085,1086,157,1002],"AI","H200","MI300X",{"path":1274,"title":1275,"description":1276,"date":1277,"slug":1278,"image":1279,"originalUrl":1280,"categories":1281},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Explore ElioVP's approach to local AI for business workflows, including custom model training and automated damage detection for logistics.","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[1058,1036,1156,1157,1255,1256,1257,1258,1162,1164,1089,1259,1260],{"path":1283,"title":1284,"description":1285,"date":1286,"slug":1287,"image":81,"originalUrl":1288,"categories":1289},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Test Paiton with free evaluation models for AMD GPUs. Compare text, vision and image generation performance using your own workloads.","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[1058,1036,157],{"path":1291,"title":1292,"description":1293,"date":1294,"slug":1295,"image":1296,"originalUrl":1297,"categories":1298},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Llama 3.1 405B: Faster Startup with Paiton on MI300X","See Paiton benchmarks for Llama 3.1 405B on eight AMD MI300X GPUs, covering model startup, tensor parallelism, throughput and latency.","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[1058,1036,157,1105,1037,1211,1299,1213,1300,1301,1302,157,1303,1304],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":1306,"title":1307,"description":1308,"date":1309,"slug":1310,"image":1311,"originalUrl":1312,"categories":1313},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","Compare Paiton on AMD MI300X with NVIDIA H200 for Llama 3.1 70B FP8, including throughput, first-token delay and latency across batch sizes.","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[1058,1036,157,1105,1211,1314,1161,1081,1038,1039,1040,1301,1315,1316],"Cold Start Optimization","Model Serving","vLLM Optimization",{"path":1318,"title":1319,"description":1320,"date":1321,"slug":1322,"image":1323,"originalUrl":1324,"categories":1325},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","Compare MI300X, H200, RX 7900 XTX and Tenstorrent n300s on Llama 3 8B with vLLM, including throughput, modeled token costs and hardware limits.","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[1058,1036,157,1156,1105,1074,1272,1086,1326,1327],"RX7900XTX","tenstorrent",{"path":1329,"title":1330,"description":1331,"date":1332,"slug":1333,"image":1334,"originalUrl":1335,"categories":1336},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Financial Modeling for GPU Clusters","Explore how ClusterP&L models GPU cluster costs, profitability and investment scenarios, with ROI metrics, risk simulations and exportable reports.","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[1058,1036,1104,1156,1075,1271,1337,1086,1338],"MI325x","pnl calculator",{"path":1340,"title":1341,"description":1342,"date":1343,"slug":1344,"image":1345,"originalUrl":1346,"categories":1347},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","AMD MI300X vs. NVIDIA H200: Qwen3-32B with Paiton","Compare Qwen3-32B benchmarks on Paiton-optimized AMD MI300X and NVIDIA H200, covering throughput, latency and hardware costs.","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[1058,1036,157,1270,1074,1271,1348,1086,157,1002],"MI300",{"path":1350,"title":1351,"description":1352,"date":1353,"slug":1354,"image":1355,"originalUrl":1356,"categories":1357},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Modular Data Centers for NVIDIA NVL: 1 to 2 MW","Explore modular data center designs for NVIDIA NVL systems, covering power capacity, liquid cooling, redundancy and deployment planning.","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[1058,1104,1358,1059,1195,1062,1196,1197,1359,1360,1200,1361],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":1363,"title":1364,"description":1365,"date":1366,"slug":1367,"image":1368,"originalUrl":1369,"categories":1370},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","Explore a local AI agent demo for DICOM workflows, from patient and study retrieval to a comparison of vision models using anonymized medical images.","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[1058,1036,1156,1105,1270,1074,1371],"Healthcare",{"path":1373,"title":1374,"description":1375,"date":1376,"slug":1377,"image":1378,"originalUrl":1379,"categories":1380},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","U.S. Tariffs and AI Supply Chain Resilience: April 2025","Read ElioVP's April 2025 perspective on U.S. tariffs and supply chain resilience for AI servers, HPC systems and modular data centers.","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[1058,1116,1270,1074,1381,1382,1383,1384],"import","Taiwan","Tariffs","Trump",{"path":1386,"title":1387,"description":1388,"date":1389,"slug":1390,"image":1391,"originalUrl":1392,"categories":1393},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","Explore AI agents for ERP, CRM, finance and customer support, with practical use cases and a path from workflow assessment to pilot and deployment.","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[1058,1036,1156,1270,1394,1395],"AI Agents","ERP",{"path":1397,"title":1398,"description":1399,"date":1400,"slug":1401,"image":1402,"originalUrl":1403,"categories":1404},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","Explore open-source AI optimization trends, from quantization and mixture-of-experts models to hardware-aware tuning, RAG and edge deployment.","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[1058,1036,1116,1405,1074,633,1086],"AI news",{"path":1407,"title":1408,"description":1409,"date":1410,"slug":1411,"image":1412,"originalUrl":1413,"categories":1414},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","Explore our dstack-powered tool for reproducible vLLM benchmarks, automated parameter sweeps and performance reports across local and cloud GPUs.","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[1058,1036,157,1270,1074,1415,1416,1272,157],"benchmark","LLM",{"path":1418,"title":1419,"description":1420,"date":1421,"slug":1422,"image":1423,"originalUrl":1424,"categories":1425},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","Compare QwQ-32B throughput and latency on AMD MI300X with Paiton and NVIDIA H200, from small batches to higher concurrency.","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[1058,1036,157],{"path":1427,"title":1428,"description":1429,"date":1430,"slug":1431,"image":1432,"originalUrl":1433,"categories":1434},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","Listen to Elio Van Puyvelde and Jim Greene on AMD's Tech Talk podcast, discussing ElioVP's origins and its AI hardware and software services.","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[1058,1074,1435,1436,1437],"Jim Greene","Podcast","Tech Talk",{"path":1439,"title":1440,"description":1441,"date":1442,"slug":1443,"image":1444,"originalUrl":1445,"categories":1446},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Explore Paiton's DeepSeek R1 Distill Llama 8B benchmarks on AMD MI300X, focusing on throughput and first-token latency at smaller batch sizes.","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[1058,1036,157,1074,1447,1448,1271,1272,1337,157,1002],"Deepseek","H100",{"path":1450,"title":1451,"description":1452,"date":1453,"slug":1454,"image":1455,"originalUrl":1456,"categories":1457},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","Paiton Benchmarks: DeepSeek R1 Distill Llama 3.1 8B","Compare stock and Paiton-optimized DeepSeek R1 Distill Llama 3.1 8B on AMD MI300X, with throughput and latency benchmarks across batch sizes.","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[1058,1036,157,1074,1447,1448,1271,1272,1337,157,1002],{"path":1459,"title":1460,"description":1461,"date":1462,"slug":1463,"image":1464,"originalUrl":1465,"categories":1466},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","Learn how Paiton uses model compilation, custom kernels and kernel fusion to optimize AI inference on AMD GPUs.","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[1058,1036,157,1074,1448,1271,1272,1337,157,1002],1789853166179]