[{"data":1,"prerenderedAt":1234},["ShallowReactive",2],{"blog-post-en-\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700":3,"blog-posts-sidebar-en":757},{"id":4,"title":5,"body":6,"categories":738,"date":744,"description":745,"extension":746,"heading":747,"image":748,"meta":749,"navigation":32,"originalUrl":750,"path":751,"seo":752,"slug":753,"stem":754,"updated":755,"__hash__":756},"blog\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700.md","MiniMax H3 on Radeon: 15-Second Video With Native Sound",{"type":7,"value":8,"toc":728},"minimark",[9,13,24,27,39,46,52,57,72,135,142,149,157,161,168,215,232,238,241,250,254,270,339,342,353,357,364,371,378,381,387,391,394,444,447,450,456,463,469,478,482,489,495,569,572,575,587,591,598,652,662,668,675,682,693,697,700,712,721,724],[10,11,12],"p",{},"A fox steps out of the trees and pauses beside a stream. The camera follows the scene, accompanied by birds and flowing water. This is one continuous MiniMax H3 generation, with its original stereo soundtrack, created locally on a single Radeon AI PRO R9700.",[10,14,15,19,20,23],{},[16,17,18],"strong",{},"Paiton produces the playable clip in 5 minutes 33 seconds on average."," Matched stock takes 6 minutes 39 seconds. That is about ",[16,21,22],{},"67 seconds less waiting per clip",", with identical MP4 hashes in the retained comparison.",[10,25,26],{},"This is the next free Paiton RDNA community package: local video with sound, an included ComfyUI workflow, and a benchmark that runs all the way from a fresh prompt to a saved MP4. No cloud inference service is required after setup.",[10,28,29],{},[30,31],"video",{"controls":32,"playsInline":32,"preload":33,"width":34,"height":35,"ariaLabel":36,"poster":37,"src":38},true,"none",864,480,"Generated fox scene with native stereo sound, 15 seconds","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002Ffox-15s-poster.jpg","\u002Fasset\u002Fvideos\u002Fblog\u002Fpaiton-minimax-h3\u002Ffox-15s.mp4",[10,40,41,45],{},[42,43,44],"a",{"href":38},"Open the fox video (MP4, 5.5 MB)",". Press play to watch and hear the original sound.",[10,47,48],{},[49,50,51],"em",{},"One generated scene: 864 × 480, 362 frames at 24 fps, or 15.0833 seconds of video. Native 32 kHz stereo audio. No frame interpolation, upscaling, soundtrack replacement or clip concatenation. The 15 seconds describe the output, not the time needed to generate it.",[53,54,56],"h2",{"id":55},"the-same-clip-about-a-minute-sooner","The same clip, about a minute sooner",[10,58,59,60,63,64,67,68,71],{},"At the matched ",[16,61,62],{},"Turbo8 profile with eight denoiser evaluations",", Paiton reduces complete-request latency by ",[16,65,66],{},"16.66%",". At that generation rate, the same active runtime projects to ",[16,69,70],{},"19.99% more clips per hour",".",[73,74,75,92],"table",{},[76,77,78],"thead",{},[79,80,81,85,89],"tr",{},[82,83,84],"th",{},"Complete request, matched Turbo8",[82,86,88],{"align":87},"right","Stock",[82,90,91],{"align":87},"Paiton",[93,94,95,109,122],"tbody",{},[79,96,97,101,104],{},[98,99,100],"td",{},"Mean time to a saved, playable MP4",[98,102,103],{"align":87},"399.40 s",[98,105,106],{"align":87},[16,107,108],{},"332.85 s",[79,110,111,114,117],{},[98,112,113],{},"Approximate waiting time",[98,115,116],{"align":87},"6m 39s",[98,118,119],{"align":87},[16,120,121],{},"5m 33s",[79,123,124,127,130],{},[98,125,126],{},"Projected fixed-length clips per hour",[98,128,129],{"align":87},"9.01",[98,131,132],{"align":87},[16,133,134],{},"10.82",[10,136,137],{},[138,139],"img",{"alt":140,"src":141},"MiniMax H3 complete-request latency: 399.40 seconds for stock and 332.85 seconds for Paiton. Projected throughput rises from 9.01 to 10.82 clips per hour.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002Fgeneration-performance.webp",[10,143,144,145,148],{},"The comparison uses ",[16,146,147],{},"one fox prompt, seed 771, one warmup and two measured requests per engine",". Every measured request includes fresh conditioning, denoising, video and audio decoding, H.264\u002FAAC encoding, muxing and writing the MP4. Model downloads, process setup and the initial warmup are excluded.",[10,150,151,152,156],{},"The headline is therefore not a denoising-only speedup. It measures the wait for a playable file. Clips per hour is calculated as ",[153,154,155],"code",{},"3600 \u002F mean request seconds",", not measured in an hour-long production run. Prompt writing, review and rejected outputs add time in actual use.",[53,158,160],{"id":159},"start-creating-in-comfyui","Start creating in ComfyUI",[10,162,163,164,167],{},"On a Linux workstation with a ",[16,165,166],{},"32 GB Radeon AI PRO R9700",", Docker, Compose and working Radeon device access, run:",[169,170,175],"pre",{"className":171,"code":172,"language":173,"meta":174,"style":174},"language-bash shiki shiki-themes github-light github-dark","git clone --depth 1 https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\ncd paiton-vllm-plugin\n.\u002Fmodels\u002FMiniMax-H3\u002Flaunch.sh\n","bash","",[153,176,177,200,209],{"__ignoreMap":174},[178,179,182,186,190,194,197],"span",{"class":180,"line":181},"line",1,[178,183,185],{"class":184},"sScJk","git",[178,187,189],{"class":188},"sZZnC"," clone",[178,191,193],{"class":192},"sj4cs"," --depth",[178,195,196],{"class":192}," 1",[178,198,199],{"class":188}," https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin.git\n",[178,201,203,206],{"class":180,"line":202},2,[178,204,205],{"class":192},"cd",[178,207,208],{"class":188}," paiton-vllm-plugin\n",[178,210,212],{"class":180,"line":211},3,[178,213,214],{"class":184},".\u002Fmodels\u002FMiniMax-H3\u002Flaunch.sh\n",[10,216,217,218,224,225,228,229,71],{},"Open ",[42,219,223],{"href":220,"rel":221},"http:\u002F\u002F127.0.0.1:8190\u002F?paiton=1",[222],"nofollow","ComfyUI on localhost",". The included workflow opens on the first visit. Edit the prompt, choose a seed and click ",[16,226,227],{},"Run",". The engine selector offers stock and Paiton, and videos are saved in ",[153,230,231],{},"~\u002Fpaiton-videos\u002F",[10,233,234],{},[138,235],{"alt":236,"src":237},"The supplied ComfyUI workflow, including the MiniMax H3 engine selector, Turbo8 settings, prompt, seed, and video and audio decoding.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002Fcomfyui-workflow.webp",[10,239,240],{},"The first launch downloads the artifact image and SHA-256-verified checkpoints, and installs pinned ComfyUI components locally. Later launches reuse the models, image and caches. Initial setup needs internet access; cached generation does not need a model API, paid inference service or cloud GPU.",[10,242,243,244,249],{},"The ",[42,245,248],{"href":246,"rel":247},"https:\u002F\u002Fgithub.com\u002FEliovp-BV\u002Fpaiton-vllm-plugin\u002Ftree\u002Fmain\u002Fmodels\u002FMiniMax-H3",[222],"model guide"," contains the setup instructions, terminal commands, supported settings and license terms. Review the model and encoder licenses before use; the runtime does not replace those terms. The command above follows the repository's current default branch; use a recorded release or commit when reproducing a specific benchmark.",[53,251,253],{"id":252},"what-your-workstation-needs","What your workstation needs",[10,255,256,257,260,261,265,266,269],{},"This is a ",[16,258,259],{},"32 GB GPU workload",", not the lower-memory image-generation profile from our ",[42,262,264],{"href":263},"\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700","FLUX.2 klein release",". The long-clip record reaches ",[16,267,268],{},"30.34 GiB of sampled driver memory",". Components are scheduled with memory management between phases; the complete pipeline is not claimed to remain in VRAM at once.",[73,271,272,282],{},[76,273,274],{},[79,275,276,279],{},[82,277,278],{},"Requirement",[82,280,281],{},"Tested configuration or recommendation",[93,283,284,292,300,311,323,331],{},[79,285,286,289],{},[98,287,288],{},"GPU",[98,290,291],{},"One AMD Radeon AI PRO R9700, 32 GB",[79,293,294,297],{},[98,295,296],{},"Host",[98,298,299],{},"Intel i5-8400, 16 GB system RAM and 4 GB swap",[79,301,302,305],{},[98,303,304],{},"Recommended system RAM",[98,306,307,310],{},[16,308,309],{},"24 GB or more"," for the browser, longer requests and other applications",[79,312,313,316],{},[98,314,315],{},"Storage",[98,317,318,319,322],{},"SSD with ",[16,320,321],{},"60 GB free"," for the selected setup, caches and outputs",[79,324,325,328],{},[98,326,327],{},"Selected checkpoints",[98,329,330],{},"Approximately 33.96 GB; both adapters and their shared components total 35.92 GB",[79,332,333,336],{},[98,334,335],{},"Filesystem consideration",[98,337,338],{},"Without hard-link support, duplicate cache copies can require another 34 GB",[10,340,341],{},"The 16 GB host used swap. Across the retained records, lifetime process RSS reached 12.09 GiB and sampled process swap reached 1.24 GiB. Those figures are not evidence that every 16 GB system will run comfortably. Driver-memory telemetry was sampled every 0.5 seconds and may miss brief peaks.",[10,343,344,345,348,349,352],{},"With weights already present, the first 15-second requests took ",[16,346,347],{},"423.6 seconds for stock"," and ",[16,350,351],{},"381.0 seconds for Paiton",", after process setup. Downloads, verification, image preparation and cold caches add further time. The 332.85-second result is the mean of the two measured Paiton requests after that warmup, not first-launch time.",[53,354,356],{"id":355},"same-settings-same-retained-output","Same settings, same retained output",[10,358,359,360,363],{},"Both engines use the same upstream pruned and quantized ",[16,361,362],{},"W4A8 generation profile",", Turbo adapter, prompt, seed, frame count, scheduler and guidance. Paiton optimizes execution of that profile. It does not claim the upstream model pruning, quantization or Turbo training as its own work.",[10,365,366,367,370],{},"For the retained 15-second fox case, ",[16,368,369],{},"all six recorded MP4 outputs have the same SHA-256 hash",": a warmup and two measured requests from each engine. The complete results record and per-run CSV are included below. This supports exact output agreement for this test, not a promise of identical files for every model, prompt or future runtime.",[10,372,373,374,377],{},"The original full-precision 33B model is ",[16,375,376],{},"not"," the quality baseline. Both measured engines already include the upstream compression and adapter choices. MiniMax's hosted context-processing system and 2K regeneration stage are also outside this package. This is the local H3 base-generation path, not a claim to reproduce the entire hosted service.",[10,379,380],{},"The supplied quality review describes a consistent fox across the sequence, distinct decoded frames and natural sound. That is a useful retained example, not a comprehensive video-quality evaluation.",[10,382,383],{},[138,384],{"alt":385,"src":386},"Six retained frames from the fox clip at approximately 0, 3, 6, 9, 12 and 15 seconds.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002Ffox-15s-contact-sheet.jpg",[53,388,390],{"id":389},"four-steps-are-an-option-not-a-hidden-benchmark-shortcut","Four steps are an option, not a hidden benchmark shortcut",[10,392,393],{},"An earlier short-clip suite covered wildlife, a speaking barista and pouring water. These outputs contain 124 frames at 24 fps, or approximately 5.17 seconds of video, with native stereo audio.",[73,395,396,410],{},[76,397,398],{},[79,399,400,403,405,407],{},[82,401,402],{},"Earlier short-clip suite",[82,404,88],{"align":87},[82,406,91],{"align":87},[82,408,409],{"align":87},"Lower latency",[93,411,412,428],{},[79,413,414,417,420,425],{},[98,415,416],{},"Turbo8, eight evaluations",[98,418,419],{"align":87},"95.83 s",[98,421,422],{"align":87},[16,423,424],{},"80.60 s",[98,426,427],{"align":87},"15.89%",[79,429,430,433,436,441],{},[98,431,432],{},"Turbo4, four evaluations",[98,434,435],{"align":87},"64.15 s",[98,437,438],{"align":87},[16,439,440],{},"54.38 s",[98,442,443],{"align":87},"15.24%",[10,445,446],{},"These are separate short-clip results reported in the release notes, not additional samples in the 15-second benchmark. Each row compares the same adapter and evaluation count across engines.",[10,448,449],{},"The four-step adapter gives a faster iteration option, but it changes the quality tradeoff. In the pouring test, it produced two bottles where one was requested. Eight steps preserved one bottle but still showed excessive foam and incomplete placement. Both versions had a brief native audio transient.",[10,451,452,455],{},[16,453,454],{},"Turbo8 remains the quality-oriented default."," We do not compare four-step Paiton against eight-step stock and call the reduced work a compiler speedup.",[10,457,458],{},[30,459],{"controls":32,"playsInline":32,"preload":33,"width":34,"height":35,"ariaLabel":460,"poster":461,"src":462},"Generated barista dialogue with native sound, 5 seconds","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002Fbarista-5s-poster.jpg","\u002Fasset\u002Fvideos\u002Fblog\u002Fpaiton-minimax-h3\u002Fbarista-5s.mp4",[10,464,465,468],{},[42,466,467],{"href":462},"Open the barista video (MP4, 0.7 MB)",". Spoken line: “Here is your coffee.”",[10,470,471],{},[49,472,473,474,71],{},"Turbo4 dialogue example, approximately 5.17 seconds. The supplied reviewer notes report intelligible “Here is your coffee” speech and reasonable lip synchronization on the byte-identical stock counterpart. This is a separate example from the timed 15-second fox comparison. ",[42,475,477],{"href":476},"\u002Fdownloads\u002Fpaiton-minimax-h3\u002Fbarista-provenance.json","Clip provenance",[53,479,481],{"id":480},"measure-the-whole-request","Measure the whole request",[10,483,484,485,488],{},"The benchmark ran on the workstation's existing ",[16,486,487],{},"AUTO\u002FCOMPUTE profile",". We did not alter clock, power, voltage or cooling controls. Conditioning is recomputed for each request rather than reused from a previous prompt.",[10,490,491],{},[138,492],{"alt":493,"src":494},"Mean time spent on conditioning, denoising, video decoding, audio decoding, and encoding and muxing for stock and Paiton.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002Fcomponent-times.webp",[73,496,497,508],{},[76,498,499],{},[79,500,501,504,506],{},[82,502,503],{},"Mean time per request phase",[82,505,88],{"align":87},[82,507,91],{"align":87},[93,509,510,523,536,547,558],{},[79,511,512,515,518],{},[98,513,514],{},"Conditioning",[98,516,517],{"align":87},"21.72 s",[98,519,520],{"align":87},[16,521,522],{},"10.82 s",[79,524,525,528,531],{},[98,526,527],{},"Denoising",[98,529,530],{"align":87},"317.99 s",[98,532,533],{"align":87},[16,534,535],{},"262.30 s",[79,537,538,541,544],{},[98,539,540],{},"Video decoding",[98,542,543],{"align":87},"42.52 s",[98,545,546],{"align":87},"42.47 s",[79,548,549,552,555],{},[98,550,551],{},"Audio decoding",[98,553,554],{"align":87},"3.35 s",[98,556,557],{"align":87},"3.37 s",[79,559,560,563,566],{},[98,561,562],{},"Encoding and muxing",[98,564,565],{"align":87},"13.80 s",[98,567,568],{"align":87},"13.87 s",[10,570,571],{},"The phase measurements account for essentially the complete request. Small bookkeeping intervals sit outside those phases. Keeping the saved-file boundary matters: accelerating model execution does not make video decoding, audio decoding or file encoding disappear.",[10,573,574],{},"This is a small, fixed-seed comparison on one workstation. It establishes the result for the published profile, not a speed guarantee for all prompts or a claim to be faster than every other H3 runtime.",[10,576,577,578,348,582,586],{},"The downloadable ",[42,579,581],{"href":580},"\u002Fdownloads\u002Fpaiton-minimax-h3\u002Flong15-timings.csv","per-run timings",[42,583,585],{"href":584},"\u002Fdownloads\u002Fpaiton-minimax-h3\u002Fbenchmark-results.json","benchmark results"," preserve the measured values, settings, output hashes and memory summaries. Paiton's compiler implementation remains separate from this public benchmark evidence.",[53,588,590],{"id":589},"lower-cost-per-generated-minute","Lower cost per generated minute",[10,592,593,594,597],{},"A video workload calls for video economics. We count ",[16,595,596],{},"cost per complete clip and per generated minute",", not internal diffusion tokens.",[73,599,600,611],{},[76,601,602],{},[79,603,604,607,609],{},[82,605,606],{},"Illustrative ownership cost",[82,608,88],{"align":87},[82,610,91],{"align":87},[93,612,613,626,639],{},[79,614,615,618,621],{},[98,616,617],{},"USD per 15.08-second clip",[98,619,620],{"align":87},"$0.0373",[98,622,623],{"align":87},[16,624,625],{},"$0.0311",[79,627,628,631,634],{},[98,629,630],{},"USD per generated video minute",[98,632,633],{"align":87},"$0.148",[98,635,636],{"align":87},[16,637,638],{},"$0.124",[79,640,641,644,647],{},[98,642,643],{},"EUR per generated video minute",[98,645,646],{"align":87},"€0.205",[98,648,649],{"align":87},[16,650,651],{},"€0.171",[10,653,654,655,658,659,71],{},"At equal assumed cost per productive hour, the measured latency reduction translates into ",[16,656,657],{},"16.66% lower modeled cost per generated minute",". The inverse is ",[16,660,661],{},"19.99% more generated video for the same active-runtime budget",[10,663,664],{},[138,665],{"alt":666,"src":667},"Modeled cost per generated video minute falls by 16.7% at equal hourly cost. The US scenario moves from $0.148 to $0.124; the European scenario from €0.205 to €0.171.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002Fmodeled-cost.webp",[10,669,670,671,674],{},"The scenarios amortize a $1,299 GPU with $0.17\u002FkWh electricity, or a €1,749 GPU with €0.2558\u002FkWh electricity, over 5,000 productive hours. Both engines are assigned an equal ",[16,672,673],{},"450 W whole-system draw",". These are scenario inputs, not current retail quotes or measured wall power.",[10,676,677,678,681],{},"The calculation excludes the host purchase, idle time, cooling, maintenance, financing, tax, residual value, manual review and rejected generations. ",[16,679,680],{},"If only half of the generated clips are usable, cost per accepted clip doubles."," These are generation costs, not the price of a finished production minute.",[10,683,243,684,348,688,692],{},[42,685,687],{"href":686},"\u002Fdownloads\u002Fpaiton-minimax-h3\u002Feconomics.json","scenario data",[42,689,691],{"href":690},"\u002Fdownloads\u002Fpaiton-minimax-h3\u002Feconomics.py","calculator"," are included so you can substitute your own hardware cost, utilization and electricity rate.",[53,694,696],{"id":695},"try-the-local-package-bring-us-the-production-workload","Try the local package. Bring us the production workload.",[10,698,699],{},"This release adds video with native sound to Paiton's free RDNA community packages, alongside local language models and image generation. The immediate result is practical: the same retained clip, on the same Radeon, saved about a minute sooner.",[10,701,702,706,707,71],{},[42,703,705],{"href":246,"rel":704},[222],"Get the MiniMax H3 setup and benchmarks",", or inspect the ",[42,708,711],{"href":709,"rel":710},"https:\u002F\u002Fhuggingface.co\u002FEliovpAI\u002FMiniMax-H3-W4A8-Paiton-RDNA4",[222],"Hugging Face runtime package",[10,713,714,717,718,71],{},[16,715,716],{},"Running production inference on AMD Instinct or CDNA?"," Talk to us about your models, latency targets and cost per usable result. This Radeon benchmark is one qualified workload; larger deployments need their own measurements. Learn more about ",[42,719,91],{"href":720},"\u002Fproducts\u002Fpaiton",[10,722,723],{},"Powered by MiniMax H3.",[725,726,727],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":174,"searchDepth":202,"depth":202,"links":729},[730,731,732,733,734,735,736,737],{"id":55,"depth":202,"text":56},{"id":159,"depth":202,"text":160},{"id":252,"depth":202,"text":253},{"id":355,"depth":202,"text":356},{"id":389,"depth":202,"text":390},{"id":480,"depth":202,"text":481},{"id":589,"depth":202,"text":590},{"id":695,"depth":202,"text":696},[91,739,740,741,742,743],"AMD Radeon","Local AI","Video Generation","MiniMax H3","ComfyUI","2026-09-09T07:30:00Z","Paiton generates a 15-second MiniMax H3 video with stereo audio on one Radeon AI PRO R9700 in 5m 33s, with 16.7% lower latency than matched stock.","md","15 seconds of video. Native sound. One Radeon.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002F00-featured-minimax-h3-r9700.webp",{},"https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700","\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700",{"title":5,"description":745},"paiton-minimax-h3-radeon-ai-pro-r9700","blog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700",null,"e7k59G4uu-RyDj5bu5gg8kArL62JouW3gvR1pH6OJC4",[758,771,781,783,792,807,816,830,861,873,895,913,932,950,968,985,997,1013,1028,1040,1049,1057,1072,1084,1095,1106,1116,1129,1139,1152,1163,1173,1184,1193,1205,1216,1225],{"path":759,"title":760,"description":761,"date":762,"slug":763,"image":764,"originalUrl":765,"categories":766},"\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700","Qwen3.8: 400.7 tok\u002Fs on R9700 | Paiton","Qwen3.8 on one R9700: 400.7 aggregate tok\u002Fs with ROCm 10 and vLLM 0.29, plus public 200K\u002F220K chat profiles. Benchmarks, limits and launch commands.","2026-09-16T07:30:00Z","paiton-qwen38-mxfp4-dflash2-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002Fupdate-2026-09-19\u002Fhero.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700",[91,739,767,768,769,770],"vLLM","Qwen3.8","Inference Optimization","DFlash2",{"path":772,"title":773,"description":774,"date":775,"slug":776,"image":777,"originalUrl":778,"categories":779},"\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700","Qwen3.8 GGUF in vLLM: Faster Responses on One Radeon","Run the original NEO CODER MAX GGUF in vLLM with Paiton on an R9700. Explore measured latency gains, image input and local deployment.","2026-09-14T07:30:00Z","paiton-qwen38-neo-gguf-vllm-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-neo-gguf\u002F00-hero-neo-gguf-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700",[91,739,740,780,767],"GGUF",{"path":751,"title":5,"description":745,"date":744,"slug":753,"image":748,"originalUrl":750,"categories":782},[91,739,740,741,742,743],{"path":263,"title":784,"description":785,"date":786,"slug":787,"image":788,"originalUrl":263,"categories":789},"Local FLUX.2 klein on Radeon AI PRO R9700: Faster Image Generation with Less VRAM","Paiton generates 1024×1024 FLUX.2 klein images in 1.054 seconds on an R9700: 16.2% lower warm latency and 33.4% lower peak Torch allocation.","2026-09-07T09:00:00","paiton-flux2-klein-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Ffox-paiton.webp",[91,739,740,790,791,743],"Image Generation","FLUX",{"path":793,"title":794,"description":795,"date":796,"slug":797,"image":798,"originalUrl":799,"categories":800},"\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700","Ornith 1.5 at 44.6 tok\u002Fs on One Radeon AI PRO R9700","Paiton serves Ornith 1.5 35B A3B at 44.63 output tok\u002Fs on one Radeon AI PRO R9700, 27% faster than tuned stock vLLM, with 21.3% lower modeled cost.","2026-09-05T09:00:00","paiton-ornith15-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-ornith15\u002F00-featured-ornith15-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700",[91,801,739,802,803,804,769,805,767,806],"Artificial Intelligence","AI Inference","GPU Performance","Inference Latency","Large Language Models","Cost Efficiency",{"path":808,"title":809,"description":810,"date":811,"slug":812,"image":813,"originalUrl":814,"categories":815},"\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","Paiton: 21% More Qwen3.8 Throughput on Radeon AI PRO R9700","Paiton serves AMD’s Qwen3.8 27B at 39.77 output tokens\u002Fs on one Radeon AI PRO R9700, delivering 21% more throughput and 17.4% lower modeled cost.","2026-09-04T09:00:00","paiton-qwen38-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F00-featured-paiton-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700",[91,801,739,802,803,804,769,805,767,806],{"path":817,"title":818,"description":819,"date":820,"slug":821,"image":822,"originalUrl":755,"categories":823},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",[824,825,826,827,828,829],"All","AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":831,"title":832,"description":833,"date":834,"slug":835,"image":836,"originalUrl":837,"categories":838},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Wan2.2 Video Generation: Paiton on AMD MI355X","Compare Wan2.2-T2V-A14B video generation on AMD MI355X with Paiton and NVIDIA B200 using Diffusers, and explore our diffusion optimization approach.","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[824,801,91,839,840,841,842,843,844,845,846,847,288,848,849,850,851,852,853,854,91,855,856,857,858,859,860],"14B","AMD","B200","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","Hardware","Inference","Instinct","MI355x","NVidia","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":862,"title":863,"description":864,"date":865,"slug":866,"image":867,"originalUrl":868,"categories":869},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","ElioVP in De Tijd: Chip Optimization and Data Centers","Read about De Tijd's coverage of ElioVP, from its origins in chip optimization to its work on modular data centers and high-density cooling.","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[824,801,870,871,840,872,870,852],"Modular DC","Uncategorized","De Tijd",{"path":874,"title":875,"description":876,"date":877,"slug":878,"image":879,"originalUrl":880,"categories":881},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","AI Privacy: A Strategic Priority for Benelux Businesses","Explore generative AI privacy risks, trust, data retention and governance, and why Benelux businesses need a strategic approach to secure AI.","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[824,801,882,871,883,884,885,886,887,888,889,890,846,891,847,892,893,894],"Trending","AI Act","Anthropomorphism","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Microsoft Copilot","Privacy","Shadow AI",{"path":896,"title":897,"description":898,"date":899,"slug":900,"image":901,"originalUrl":902,"categories":903},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Why We Do Not Use itsme: Privacy and Data Sovereignty","Why ElioVP does not use itsme: our assessment of identity metadata, cloud dependence, data sovereignty and authentication risks.","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[824,904,905,906,888,907,908,909,891,910,911,912,893],"AWS","Belgian Mobile ID","Cloud Act","Data Sovereignty","Digital Identity","eIDAS","itsme","Liberty Global","MyGov.be",{"path":914,"title":915,"description":916,"date":917,"slug":918,"image":919,"originalUrl":920,"categories":921},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","Lessons from building on-premise AI agents in 2025 cover workflow design, observability, model training, hallucinations and GPU memory limits.","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[824,801,922,882,923,924,925,926,927,928,929,930,855,931],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":933,"title":934,"description":935,"date":936,"slug":937,"image":938,"originalUrl":939,"categories":940},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","An analysis of AI neocloud investment risks, examining circular financing, infrastructure claims, contract terms and due diligence.","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[824,801,882,825,941,942,943,944,945,946,947,948,949],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":951,"title":952,"description":953,"date":954,"slug":955,"image":956,"originalUrl":957,"categories":958},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","NVIDIA GB300 NVL72: A Four-Month Modular Data Center Plan","Explore a modular data center design for NVIDIA GB300 NVL72, covering redundant power, hybrid cooling and a four-month deployment plan.","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[824,870,871,959,825,960,961,962,963,964,965,966,967],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":969,"title":970,"description":971,"date":972,"slug":973,"image":974,"originalUrl":975,"categories":976},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Why CUDA compatibility is not the same as AMD performance: explore ROCm, HIP, kernel tuning and the case for hardware-specific optimization.","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[824,801,91,871,977,801,978,979,980,981,982,983,91,984],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","HIP","Kernel Tuning","ROCm",{"path":986,"title":987,"description":988,"date":989,"slug":990,"image":991,"originalUrl":992,"categories":993},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Learn how Paiton integrates with existing inference stacks, with AMD MI300X benchmark results and performance-per-dollar comparisons.","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[824,801,91,802,994,977,806,995,769,983,91,996,767],"AMD Instinct","High Throughput","SGLang",{"path":998,"title":999,"description":1000,"date":1001,"slug":1002,"image":1003,"originalUrl":1004,"categories":1005},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Paiton MoE Benchmarks: MI300X vs H200 and B200","Compare Qwen3-30B-A3B MoE inference with Paiton on MI300X against H200 and B200, including throughput and cost per million tokens.","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[824,801,91,1006,977,1007,769,1008,1009,1010,1011,91,1012],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":1014,"title":1015,"description":1016,"date":1017,"slug":1018,"image":1019,"originalUrl":1020,"categories":1021},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Local Agentic AI: From Inbox to Action","Local-first AI agents turn email, documents and images into tickets, reports and actions, using models tailored to your data and systems.","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[824,801,922,871,923,1022,1023,1024,1025,928,930,855,1026,1027],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":1029,"title":1030,"description":1031,"date":1032,"slug":1033,"image":1034,"originalUrl":1035,"categories":1036},"\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","MI300X FP8 Benchmarks: GPU Partitioning with Paiton","Explore Llama 3.1 8B FP8 benchmarks on partitioned MI300X GPUs with Paiton, comparing throughput and latency against NVIDIA H200 and B200.","2025-07-31T13:32:57","mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F",[824,801,91,871,1037,840,841,1038,1039,851,852,91,767],"AI","H200","MI300X",{"path":1041,"title":1042,"description":1043,"date":1044,"slug":1045,"image":1046,"originalUrl":1047,"categories":1048},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Explore ElioVP's approach to local AI for business workflows, including custom model training and automated damage detection for logistics.","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[824,801,922,923,1022,1023,1024,1025,928,930,855,1026,1027],{"path":1050,"title":1051,"description":1052,"date":1053,"slug":1054,"image":174,"originalUrl":1055,"categories":1056},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Test Paiton with free evaluation models for AMD GPUs. Compare text, vision and image generation performance using your own workloads.","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[824,801,91],{"path":1058,"title":1059,"description":1060,"date":1061,"slug":1062,"image":1063,"originalUrl":1064,"categories":1065},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Llama 3.1 405B: Faster Startup with Paiton on MI300X","See Paiton benchmarks for Llama 3.1 405B on eight AMD MI300X GPUs, covering model startup, tensor parallelism, throughput and latency.","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[824,801,91,871,802,977,1066,979,1067,1068,1069,91,1070,1071],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":1073,"title":1074,"description":1075,"date":1076,"slug":1077,"image":1078,"originalUrl":1079,"categories":1080},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","Compare Paiton on AMD MI300X with NVIDIA H200 for Llama 3.1 70B FP8, including throughput, first-token delay and latency across batch sizes.","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[824,801,91,871,977,1081,927,847,803,804,805,1068,1082,1083],"Cold Start Optimization","Model Serving","vLLM Optimization",{"path":1085,"title":1086,"description":1087,"date":1088,"slug":1089,"image":1090,"originalUrl":1091,"categories":1092},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","Compare MI300X, H200, RX 7900 XTX and Tenstorrent n300s on Llama 3 8B with vLLM, including throughput, modeled token costs and hardware limits.","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[824,801,91,922,871,840,1039,852,1093,1094],"RX7900XTX","tenstorrent",{"path":1096,"title":1097,"description":1098,"date":1099,"slug":1100,"image":1101,"originalUrl":1102,"categories":1103},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Financial Modeling for GPU Clusters","Explore how ClusterP&L models GPU cluster costs, profitability and investment scenarios, with ROI metrics, risk simulations and exportable reports.","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[824,801,870,922,841,1038,1104,852,1105],"MI325x","pnl calculator",{"path":1107,"title":1108,"description":1109,"date":1110,"slug":1111,"image":1112,"originalUrl":1113,"categories":1114},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","AMD MI300X vs. NVIDIA H200: Qwen3-32B with Paiton","Compare Qwen3-32B benchmarks on Paiton-optimized AMD MI300X and NVIDIA H200, covering throughput, latency and hardware costs.","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[824,801,91,1037,840,1038,1115,852,91,767],"MI300",{"path":1117,"title":1118,"description":1119,"date":1120,"slug":1121,"image":1122,"originalUrl":1123,"categories":1124},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Modular Data Centers for NVIDIA NVL: 1 to 2 MW","Explore modular data center designs for NVIDIA NVL systems, covering power capacity, liquid cooling, redundancy and deployment planning.","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[824,870,1125,825,961,828,962,963,1126,1127,966,1128],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":1130,"title":1131,"description":1132,"date":1133,"slug":1134,"image":1135,"originalUrl":1136,"categories":1137},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","Explore a local AI agent demo for DICOM workflows, from patient and study retrieval to a comparison of vision models using anonymized medical images.","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[824,801,922,871,1037,840,1138],"Healthcare",{"path":1140,"title":1141,"description":1142,"date":1143,"slug":1144,"image":1145,"originalUrl":1146,"categories":1147},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","U.S. Tariffs and AI Supply Chain Resilience: April 2025","Read ElioVP's April 2025 perspective on U.S. tariffs and supply chain resilience for AI servers, HPC systems and modular data centers.","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[824,882,1037,840,1148,1149,1150,1151],"import","Taiwan","Tariffs","Trump",{"path":1153,"title":1154,"description":1155,"date":1156,"slug":1157,"image":1158,"originalUrl":1159,"categories":1160},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","Explore AI agents for ERP, CRM, finance and customer support, with practical use cases and a path from workflow assessment to pilot and deployment.","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[824,801,922,1037,1161,1162],"AI Agents","ERP",{"path":1164,"title":1165,"description":1166,"date":1167,"slug":1168,"image":1169,"originalUrl":1170,"categories":1171},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","Explore open-source AI optimization trends, from quantization and mixture-of-experts models to hardware-aware tuning, RAG and edge deployment.","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[824,801,882,1172,840,288,852],"AI news",{"path":1174,"title":1175,"description":1176,"date":1177,"slug":1178,"image":1179,"originalUrl":1180,"categories":1181},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","Explore our dstack-powered tool for reproducible vLLM benchmarks, automated parameter sweeps and performance reports across local and cloud GPUs.","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[824,801,91,1037,840,1182,1183,1039,91],"benchmark","LLM",{"path":1185,"title":1186,"description":1187,"date":1188,"slug":1189,"image":1190,"originalUrl":1191,"categories":1192},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","Compare QwQ-32B throughput and latency on AMD MI300X with Paiton and NVIDIA H200, from small batches to higher concurrency.","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[824,801,91],{"path":1194,"title":1195,"description":1196,"date":1197,"slug":1198,"image":1199,"originalUrl":1200,"categories":1201},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","Listen to Elio Van Puyvelde and Jim Greene on AMD's Tech Talk podcast, discussing ElioVP's origins and its AI hardware and software services.","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[824,840,1202,1203,1204],"Jim Greene","Podcast","Tech Talk",{"path":1206,"title":1207,"description":1208,"date":1209,"slug":1210,"image":1211,"originalUrl":1212,"categories":1213},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Explore Paiton's DeepSeek R1 Distill Llama 8B benchmarks on AMD MI300X, focusing on throughput and first-token latency at smaller batch sizes.","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[824,801,91,840,1214,1215,1038,1039,1104,91,767],"Deepseek","H100",{"path":1217,"title":1218,"description":1219,"date":1220,"slug":1221,"image":1222,"originalUrl":1223,"categories":1224},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","Paiton Benchmarks: DeepSeek R1 Distill Llama 3.1 8B","Compare stock and Paiton-optimized DeepSeek R1 Distill Llama 3.1 8B on AMD MI300X, with throughput and latency benchmarks across batch sizes.","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[824,801,91,840,1214,1215,1038,1039,1104,91,767],{"path":1226,"title":1227,"description":1228,"date":1229,"slug":1230,"image":1231,"originalUrl":1232,"categories":1233},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","Learn how Paiton uses model compilation, custom kernels and kernel fusion to optimize AI inference on AMD GPUs.","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[824,801,91,840,1215,1038,1039,1104,91,767],1789853166432]