[{"data":1,"prerenderedAt":1442},["ShallowReactive",2],{"blog-post-en-\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700":3,"blog-posts-sidebar-en":969},{"id":4,"title":5,"body":6,"categories":946,"date":956,"description":957,"extension":958,"heading":959,"image":960,"meta":961,"navigation":962,"originalUrl":963,"path":964,"seo":965,"slug":966,"stem":967,"updated":959,"__hash__":968},"blog\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700.md","Paiton: 21% More Qwen3.8 Throughput on Radeon AI PRO R9700",{"type":7,"value":8,"toc":933},"minimark",[9,16,31,38,49,56,61,64,157,160,163,167,170,173,195,198,201,207,211,214,221,224,273,340,347,350,356,359,363,366,369,372,375,379,382,385,388,392,395,398,401,405,408,517,520,538,541,561,576,586,590,593,768,771,775,885,888,891,895,910,913,917,920,923,929],[10,11,12],"p",{},[13,14,15],"strong",{},"One GPU. The same public source checkpoint. The same benchmark requests. More output.",[10,17,18,19,26,27,30],{},"Paiton served AMD’s public\n",[20,21,25],"a",{"href":22,"rel":23},"https:\u002F\u002Fhuggingface.co\u002Famd\u002FQwen3.8-27B-Quark-Qronos-INT4-W4A16",[24],"nofollow","Qwen3.8-27B-Quark-Qronos-INT4-W4A16","\nat ",[13,28,29],{},"39.77 output tokens per second"," in our batch-one interactive benchmark on a\nsingle 32 GB AMD Radeon AI PRO R9700.",[10,32,33,34,37],{},"Our fastest qualified stock vLLM configuration on the same system reached\n32.86 output tokens per second. That gives Paiton ",[13,35,36],{},"21.0% more output per active\ninference hour"," from the same GPU.",[10,39,40,41,44,45,48],{},"The advantage widened to ",[13,42,43],{},"54.3%"," in the coding workload and reached ",[13,46,47],{},"26.8%","\nwith a 4,096-token input.",[10,50,51],{},[52,53],"img",{"alt":54,"src":55},"Paiton delivers 39.77 output tokens per second on the interactive workload, 37.93 on coding and 25.45 on long context, outperforming the qualified stock vLLM baseline in all three tests.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F01-throughput-by-workload-eliovp.webp",[57,58,60],"h2",{"id":59},"throughput-across-three-workloads","Throughput across three workloads",[10,62,63],{},"These are single-user serving tests with one request at a time, one active\nmodel sequence, temperature zero, thinking disabled and fixed random inputs.",[65,66,67,90],"table",{},[68,69,70],"thead",{},[71,72,73,77,81,84,87],"tr",{},[74,75,76],"th",{},"Workload",[74,78,80],{"align":79},"right","Requested input \u002F output",[74,82,83],{"align":79},"Qualified stock vLLM",[74,85,86],{"align":79},"Paiton",[74,88,89],{"align":79},"Paiton advantage",[91,92,93,115,136],"tbody",{},[71,94,95,99,102,105,110],{},[96,97,98],"td",{},"Interactive",[96,100,101],{"align":79},"256 \u002F 256",[96,103,104],{"align":79},"32.86 tok\u002Fs",[96,106,107],{"align":79},[13,108,109],{},"39.77 tok\u002Fs",[96,111,112],{"align":79},[13,113,114],{},"+21.0%",[71,116,117,120,123,126,131],{},[96,118,119],{},"Coding",[96,121,122],{"align":79},"1,024 \u002F 512",[96,124,125],{"align":79},"24.58 tok\u002Fs",[96,127,128],{"align":79},[13,129,130],{},"37.93 tok\u002Fs",[96,132,133],{"align":79},[13,134,135],{},"+54.3%",[71,137,138,141,144,147,152],{},[96,139,140],{},"Long context",[96,142,143],{"align":79},"4,096 \u002F 256",[96,145,146],{"align":79},"20.07 tok\u002Fs",[96,148,149],{"align":79},[13,150,151],{},"25.45 tok\u002Fs",[96,153,154],{"align":79},[13,155,156],{},"+26.8%",[10,158,159],{},"Each workload was run twice from a fresh server. Every run used three warmups\nfollowed by 12 measured requests, and the table reports the mean of both runs.\nAll six Paiton runs completed 12 of 12 requests with zero failures and returned\nthe full requested output length.",[10,161,162],{},"Chat templating increased the actual prompt lengths to 268–270, 1,036–1,038\nand 4,108–4,110 tokens respectively.",[57,164,166],{"id":165},"latency-faster-streaming-mixed-first-token-results","Latency: faster streaming, mixed first-token results",[10,168,169],{},"Output throughput is only part of the user experience, so we are publishing\nfirst-token and streaming latency as well.",[10,171,172],{},"Paiton reduced median time per output token in every workload:",[174,175,176,183,189],"ul",{},[177,178,179,180],"li",{},"Interactive: ",[13,181,182],{},"30 ms to 24 ms",[177,184,185,186],{},"Coding: ",[13,187,188],{},"39 ms to 25 ms",[177,190,191,192],{},"Long context: ",[13,193,194],{},"37 ms to 28 ms",[10,196,197],{},"Time to first token was workload-dependent. Median interactive TTFT increased\nfrom 258 ms to 290 ms. Coding improved slightly from 757 ms to 737 ms, while\nlong-context TTFT fell from 3,337 ms to 2,984 ms.",[10,199,200],{},"The result is a clear streaming-throughput gain, not a claim that every latency\nmetric improves in every scenario.",[10,202,203],{},[52,204],{"alt":205,"src":206},"Paiton lowers time per output token across all three workloads. Time to first token is higher for the interactive test, slightly lower for coding and lower for the long-context test.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F02-latency-snapshot-eliovp.webp",[57,208,210],{"id":209},"what-the-gain-means-for-cost","What the gain means for cost",[10,212,213],{},"For owned inference hardware, throughput determines how much output a fixed\nhour of GPU time can produce. At equal hourly system cost, a 21.0% throughput\ngain gives 21.0% more tokens for the same active runtime budget.",[10,215,216,217,220],{},"Because cost per token is the inverse of throughput, the corresponding modeled\nreduction in time-based cost per million output tokens is ",[13,218,219],{},"17.4%",".",[10,222,223],{},"The examples below amortize the GPU over 5,000 productive inference hours and\nassume an equal 450 W whole-system draw for both runtimes.",[65,225,226,239],{},[68,227,228],{},[71,229,230,233,236],{},[74,231,232],{},"Assumption",[74,234,235],{"align":79},"US example",[74,237,238],{"align":79},"European example",[91,240,241,252,263],{},[71,242,243,246,249],{},[96,244,245],{},"GPU purchase price",[96,247,248],{"align":79},"$1,299",[96,250,251],{"align":79},"€1,749",[71,253,254,257,260],{},[96,255,256],{},"Electricity",[96,258,259],{"align":79},"$0.17\u002FkWh",[96,261,262],{"align":79},"€0.2558\u002FkWh",[71,264,265,268,271],{},[96,266,267],{},"Productive inference lifetime",[96,269,270],{"align":79},"5,000 hours",[96,272,270],{"align":79},[65,274,275,286],{},[68,276,277],{},[71,278,279,282,284],{},[74,280,281],{},"Interactive token economics",[74,283,83],{"align":79},[74,285,86],{"align":79},[91,287,288,301,314,327],{},[71,289,290,293,296],{},[96,291,292],{},"Hours per million output tokens",[96,294,295],{"align":79},"8.45",[96,297,298],{"align":79},[13,299,300],{},"6.98",[71,302,303,306,309],{},[96,304,305],{},"Modeled cost per million, US example",[96,307,308],{"align":79},"$2.84",[96,310,311],{"align":79},[13,312,313],{},"$2.35",[71,315,316,319,322],{},[96,317,318],{},"Modeled cost per million, European example",[96,320,321],{"align":79},"€3.93",[96,323,324],{"align":79},[13,325,326],{},"€3.25",[71,328,329,332,335],{},[96,330,331],{},"Output over 5,000 active hours",[96,333,334],{"align":79},"591.5 million",[96,336,337],{"align":79},[13,338,339],{},"715.8 million",[10,341,342,343,346],{},"At the measured interactive rates, the same card produces roughly ",[13,344,345],{},"124 million\nadditional output tokens"," over 5,000 productive hours.",[10,348,349],{},"For the same modeled $100 ownership-and-electricity budget, Paiton produces\nabout 42.6 million tokens instead of 35.2 million. In the European example,\n€100 produces about 30.8 million tokens instead of 25.4 million.",[10,351,352],{},[52,353],{"alt":354,"src":355},"At 5,000 productive inference hours, the modeled European cost is €3.93 per million output tokens for stock vLLM and €3.25 for Paiton.","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F03-economics-sensitivity-eliovp.webp",[10,357,358],{},"These figures are a transparent time-based ownership model, not a measured\nenergy-efficiency claim. They exclude the host computer, idle time, cooling,\nmaintenance, financing, tax and residual value. Substitute your own purchase\nprice, electricity rate and productive lifetime as needed. The 17.4% relative\nreduction remains the same when both runtimes carry the same hourly cost.",[57,360,362],{"id":361},"a-stock-baseline-worth-comparing-against","A stock baseline worth comparing against",[10,364,365],{},"We did not compare Paiton with an eager or default installation and call the\nresult an optimization win. We spent significant time qualifying the fastest\nstable stock configuration we could obtain on this system.",[10,367,368],{},"The stock baseline used compiled vLLM execution at optimization level 2,\nfull-and-piecewise HIP graph capture, stock RDNA hybrid W4A16 kernels, vLLM’s\nTriton GDN implementation, the ROCm attention backend and the R9700’s supported\nhigh clock policy.",[10,370,371],{},"Both sides used the same GPU, ROCm host, model revision, tokenizer, pinned vLLM\nrevision, request data, random seed, output lengths, concurrency and supported\nGPU clock policy. We retained the complete result files and averaged two\nfresh-server runs per workload rather than selecting one favorable terminal\nresult.",[10,373,374],{},"The stock logs confirm compiled execution and graph capture. They load no\nPaiton model artifact or Paiton compute kernel.",[57,376,378],{"id":377},"what-paiton-does","What Paiton does",[10,380,381],{},"At a high level, Paiton builds a qualified, model- and hardware-specific\nexecution path and integrates it with the serving runtime. The original AMD\ncheckpoint remains unchanged on disk.",[10,383,384],{},"Paiton assembles target-specific runtime artifacts during model loading. The\noptimized path should therefore not be assumed to produce output that is\nbit-for-bit identical to stock W4 execution.",[10,386,387],{},"That is where the implementation detail stops. The model-specific kernel,\nfusion, scheduling and runtime design are closed-source Paiton IP. What we do\npublish is the part customers can validate: the supported target, benchmark\nmethod, throughput, latency, quality gates, package identity and operating\nscope.",[57,389,391],{"id":390},"quality-checks","Quality checks",[10,393,394],{},"We ran a deterministic 12-case suite covering instruction following,\narithmetic, algebra, logic, factual recall, translation, structured JSON,\nPython, SQL, summarization, format constraints and modular reasoning.",[10,396,397],{},"All 12 cases passed. Repeated temperature-zero responses were byte-identical,\nand all recorded log probabilities were finite.",[10,399,400],{},"This is a practical quality gate for the stated local-chat scope. It is not a\nreplacement for a full academic accuracy evaluation, and it should not be read\nas a claim of quality parity across every task or prompt distribution.",[57,402,404],{"id":403},"run-the-local-chatbot","Run the local chatbot",[10,406,407],{},"The public image contains the Paiton runtime plugin and compiled runtime\nartifacts. It downloads the original 19.9 GB AMD checkpoint directly from\nHugging Face into a persistent Docker volume; Paiton does not redistribute the\nmodel weights.",[409,410,415],"pre",{"className":411,"code":412,"language":413,"meta":414,"style":414},"language-bash shiki shiki-themes github-light github-dark","docker run -d \\\n  --name paiton-qwen38 \\\n  --device \u002Fdev\u002Fkfd \\\n  --device \u002Fdev\u002Fdri \\\n  --group-add video \\\n  --ipc=host \\\n  --network host \\\n  --mount type=volume,src=paiton-qwen38-cache,dst=\u002Fmodels\u002Fcache \\\n  ghcr.io\u002Feliovp\u002Fpaiton-vllm-plugin@sha256:c56baf54aca1ad229829c1de26e8792806608e65ee9d06ee210b79cd49f70bc9\n","bash","",[416,417,418,438,449,460,470,481,489,500,511],"code",{"__ignoreMap":414},[419,420,423,427,431,435],"span",{"class":421,"line":422},"line",1,[419,424,426],{"class":425},"sScJk","docker",[419,428,430],{"class":429},"sZZnC"," run",[419,432,434],{"class":433},"sj4cs"," -d",[419,436,437],{"class":433}," \\\n",[419,439,441,444,447],{"class":421,"line":440},2,[419,442,443],{"class":433},"  --name",[419,445,446],{"class":429}," paiton-qwen38",[419,448,437],{"class":433},[419,450,452,455,458],{"class":421,"line":451},3,[419,453,454],{"class":433},"  --device",[419,456,457],{"class":429}," \u002Fdev\u002Fkfd",[419,459,437],{"class":433},[419,461,463,465,468],{"class":421,"line":462},4,[419,464,454],{"class":433},[419,466,467],{"class":429}," \u002Fdev\u002Fdri",[419,469,437],{"class":433},[419,471,473,476,479],{"class":421,"line":472},5,[419,474,475],{"class":433},"  --group-add",[419,477,478],{"class":429}," video",[419,480,437],{"class":433},[419,482,484,487],{"class":421,"line":483},6,[419,485,486],{"class":433},"  --ipc=host",[419,488,437],{"class":433},[419,490,492,495,498],{"class":421,"line":491},7,[419,493,494],{"class":433},"  --network",[419,496,497],{"class":429}," host",[419,499,437],{"class":433},[419,501,503,506,509],{"class":421,"line":502},8,[419,504,505],{"class":433},"  --mount",[419,507,508],{"class":429}," type=volume,src=paiton-qwen38-cache,dst=\u002Fmodels\u002Fcache",[419,510,437],{"class":433},[419,512,514],{"class":421,"line":513},9,[419,515,516],{"class":429},"  ghcr.io\u002Feliovp\u002Fpaiton-vllm-plugin@sha256:c56baf54aca1ad229829c1de26e8792806608e65ee9d06ee210b79cd49f70bc9\n",[10,518,519],{},"The first start downloads the checkpoint. With the weights already cached,\nmodel assembly takes roughly 10 to 12 minutes on our R9700. Follow startup with:",[409,521,523],{"className":411,"code":522,"language":413,"meta":414,"style":414},"docker logs -f paiton-qwen38\n",[416,524,525],{"__ignoreMap":414},[419,526,527,529,532,535],{"class":421,"line":422},[419,528,426],{"class":425},[419,530,531],{"class":429}," logs",[419,533,534],{"class":433}," -f",[419,536,537],{"class":429}," paiton-qwen38\n",[10,539,540],{},"Once the log reports that the application is ready, open the included streaming\nchat:",[409,542,544],{"className":411,"code":543,"language":413,"meta":414,"style":414},"docker exec -it paiton-qwen38 paiton-chat\n",[416,545,546],{"__ignoreMap":414},[419,547,548,550,553,556,558],{"class":421,"line":422},[419,549,426],{"class":425},[419,551,552],{"class":429}," exec",[419,554,555],{"class":433}," -it",[419,557,446],{"class":429},[419,559,560],{"class":429}," paiton-chat\n",[10,562,563,564,567,568,571,572,575],{},"Use ",[416,565,566],{},"\u002Freset"," to clear the conversation and ",[416,569,570],{},"\u002Fquit"," to exit. Thinking is off by\ndefault for responsive local conversation; pass ",[416,573,574],{},"--thinking"," when required.",[10,577,578,579,582,583,220],{},"The same server exposes an OpenAI-compatible endpoint at\n",[416,580,581],{},"http:\u002F\u002F127.0.0.1:8000\u002Fv1\u002Fchat\u002Fcompletions"," using model name ",[416,584,585],{},"qwen38",[57,587,589],{"id":588},"reproduce-the-interactive-benchmark","Reproduce the interactive benchmark",[10,591,592],{},"After the server is ready:",[409,594,596],{"className":411,"code":595,"language":413,"meta":414,"style":414},"docker exec paiton-qwen38 sh -lc '\nMODEL_DIR=\"$(find \u002Ftmp -maxdepth 2 -type d \\\n  -path \"\u002Ftmp\u002Fpaiton-qwen38-reused-*\u002Fmodel\" -print -quit)\"\nexec vllm bench serve \\\n  --backend openai-chat \\\n  --base-url http:\u002F\u002F127.0.0.1:8000 \\\n  --endpoint \u002Fv1\u002Fchat\u002Fcompletions \\\n  --model qwen38 \\\n  --tokenizer \"$MODEL_DIR\" \\\n  --dataset-name random \\\n  --seed 42 \\\n  --num-warmups 3 \\\n  --num-prompts 12 \\\n  --random-input-len 256 \\\n  --random-output-len 256 \\\n  --random-range-ratio 0 \\\n  --random-prefix-len 0 \\\n  --request-rate inf \\\n  --max-concurrency 1 \\\n  --temperature 0 \\\n  --ignore-eos \\\n  --extra-body '\\''{\"chat_template_kwargs\":{\"enable_thinking\":false}}'\\'' \\\n  --percentile-metrics ttft,tpot,itl,e2el \\\n  --metric-percentiles 50,90,95,99 \\\n  --disable-tqdm\n'\n",[416,597,598,615,620,625,630,635,640,645,650,655,661,667,673,679,685,691,697,703,709,715,721,727,744,750,756,762],{"__ignoreMap":414},[419,599,600,602,604,606,609,612],{"class":421,"line":422},[419,601,426],{"class":425},[419,603,552],{"class":429},[419,605,446],{"class":429},[419,607,608],{"class":429}," sh",[419,610,611],{"class":433}," -lc",[419,613,614],{"class":429}," '\n",[419,616,617],{"class":421,"line":440},[419,618,619],{"class":429},"MODEL_DIR=\"$(find \u002Ftmp -maxdepth 2 -type d \\\n",[419,621,622],{"class":421,"line":451},[419,623,624],{"class":429},"  -path \"\u002Ftmp\u002Fpaiton-qwen38-reused-*\u002Fmodel\" -print -quit)\"\n",[419,626,627],{"class":421,"line":462},[419,628,629],{"class":429},"exec vllm bench serve \\\n",[419,631,632],{"class":421,"line":472},[419,633,634],{"class":429},"  --backend openai-chat \\\n",[419,636,637],{"class":421,"line":483},[419,638,639],{"class":429},"  --base-url http:\u002F\u002F127.0.0.1:8000 \\\n",[419,641,642],{"class":421,"line":491},[419,643,644],{"class":429},"  --endpoint \u002Fv1\u002Fchat\u002Fcompletions \\\n",[419,646,647],{"class":421,"line":502},[419,648,649],{"class":429},"  --model qwen38 \\\n",[419,651,652],{"class":421,"line":513},[419,653,654],{"class":429},"  --tokenizer \"$MODEL_DIR\" \\\n",[419,656,658],{"class":421,"line":657},10,[419,659,660],{"class":429},"  --dataset-name random \\\n",[419,662,664],{"class":421,"line":663},11,[419,665,666],{"class":429},"  --seed 42 \\\n",[419,668,670],{"class":421,"line":669},12,[419,671,672],{"class":429},"  --num-warmups 3 \\\n",[419,674,676],{"class":421,"line":675},13,[419,677,678],{"class":429},"  --num-prompts 12 \\\n",[419,680,682],{"class":421,"line":681},14,[419,683,684],{"class":429},"  --random-input-len 256 \\\n",[419,686,688],{"class":421,"line":687},15,[419,689,690],{"class":429},"  --random-output-len 256 \\\n",[419,692,694],{"class":421,"line":693},16,[419,695,696],{"class":429},"  --random-range-ratio 0 \\\n",[419,698,700],{"class":421,"line":699},17,[419,701,702],{"class":429},"  --random-prefix-len 0 \\\n",[419,704,706],{"class":421,"line":705},18,[419,707,708],{"class":429},"  --request-rate inf \\\n",[419,710,712],{"class":421,"line":711},19,[419,713,714],{"class":429},"  --max-concurrency 1 \\\n",[419,716,718],{"class":421,"line":717},20,[419,719,720],{"class":429},"  --temperature 0 \\\n",[419,722,724],{"class":421,"line":723},21,[419,725,726],{"class":429},"  --ignore-eos \\\n",[419,728,730,733,736,739,741],{"class":421,"line":729},22,[419,731,732],{"class":429},"  --extra-body '",[419,734,735],{"class":433},"\\'",[419,737,738],{"class":429},"'{\"chat_template_kwargs\":{\"enable_thinking\":false}}'",[419,740,735],{"class":433},[419,742,743],{"class":429},"' \\\n",[419,745,747],{"class":421,"line":746},23,[419,748,749],{"class":429},"  --percentile-metrics ttft,tpot,itl,e2el \\\n",[419,751,753],{"class":421,"line":752},24,[419,754,755],{"class":429},"  --metric-percentiles 50,90,95,99 \\\n",[419,757,759],{"class":421,"line":758},25,[419,760,761],{"class":429},"  --disable-tqdm\n",[419,763,765],{"class":421,"line":764},26,[419,766,767],{"class":429},"'\n",[10,769,770],{},"One run is useful as a local check, but consumer GPU power management can make a\ncold run slower. Run a short warmup first and confirm that the memory clock has\nreached its normal loaded state before recording a comparison.",[57,772,774],{"id":773},"tested-configuration","Tested configuration",[65,776,777,787],{},[68,778,779],{},[71,780,781,784],{},[74,782,783],{},"Component",[74,785,786],{},"Tested value",[91,788,789,800,808,818,826,836,844,854,862,869,877],{},[71,790,791,794],{},[96,792,793],{},"GPU",[96,795,796,797],{},"AMD Radeon AI PRO R9700, 32 GB, ",[416,798,799],{},"gfx1201",[71,801,802,805],{},[96,803,804],{},"Model",[96,806,807],{},"AMD Qwen3.8 27B Qronos W4A16 INT4",[71,809,810,813],{},[96,811,812],{},"Model revision",[96,814,815],{},[416,816,817],{},"649ca9d47a7de5364c6fcccc0c1b4f6e542e15e2",[71,819,820,823],{},[96,821,822],{},"ROCm",[96,824,825],{},"7.14",[71,827,828,831],{},[96,829,830],{},"vLLM revision",[96,832,833],{},[416,834,835],{},"39bd959b582c85e78e7e0326d49042ce7c3c07ed",[71,837,838,841],{},[96,839,840],{},"Paiton image",[96,842,843],{},"Qwen3.8 Qronos for Radeon AI PRO R9700",[71,845,846,849],{},[96,847,848],{},"Image digest",[96,850,851],{},[416,852,853],{},"sha256:c56baf54aca1ad229829c1de26e8792806608e65ee9d06ee210b79cd49f70bc9",[71,855,856,859],{},[96,857,858],{},"Tensor parallelism",[96,860,861],{},"1",[71,863,864,867],{},[96,865,866],{},"Maximum active sequences",[96,868,861],{},[71,870,871,874],{},[96,872,873],{},"Maximum context",[96,875,876],{},"8,192 tokens",[71,878,879,882],{},[96,880,881],{},"Qualified scope",[96,883,884],{},"Text-only, single-user inference",[10,886,887],{},"The package fails closed on an unsupported GPU identity, architecture, runtime\ncontract, artifact checksum, model contract or serving configuration. Multiple\nsimultaneous requests queue.",[10,889,890],{},"This article does not claim support for other GPUs, ROCm versions,\ntensor-parallel configurations, multimodal input or production continuous\nbatching.",[57,892,894],{"id":893},"the-practical-result","The practical result",[10,896,897,898,901,902,905,906,909],{},"On one Radeon AI PRO R9700, Paiton increased interactive Qwen3.8 output\nthroughput from 32.86 to 39.77 tokens per second. That translates into ",[13,899,900],{},"21.0%\nmore output per active hour",", a ",[13,903,904],{},"17.4% reduction in modeled time-based cost\nper million output tokens",", and roughly ",[13,907,908],{},"124 million additional tokens over\n5,000 productive hours"," at the measured interactive rate.",[10,911,912],{},"The coding and long-context results show that the gain is not limited to one\nprompt shape. The latency data also makes the boundary clear: streaming became\nfaster across all three workloads, while first-token latency remained\nworkload-dependent.",[57,914,916],{"id":915},"optimize-your-inference-workload-with-paiton","Optimize your inference workload with Paiton",[10,918,919],{},"This is a qualified Radeon AI PRO R9700 result for one model and one serving\nscope. Paiton’s broader commercial work also targets AMD Instinct CDNA\naccelerators for larger inference deployments, where throughput, fleet\nutilization and cost per generated token compound across the infrastructure.",[10,921,922],{},"We benchmark the real workload, identify the runtime bottleneck, build the\nqualified AMD path and measure the delivered result against an agreed baseline.",[10,924,925,926,220],{},"Learn more about ",[20,927,86],{"href":928},"\u002Fproducts\u002Fpaiton",[930,931,932],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":414,"searchDepth":440,"depth":440,"links":934},[935,936,937,938,939,940,941,942,943,944,945],{"id":59,"depth":440,"text":60},{"id":165,"depth":440,"text":166},{"id":209,"depth":440,"text":210},{"id":361,"depth":440,"text":362},{"id":377,"depth":440,"text":378},{"id":390,"depth":440,"text":391},{"id":403,"depth":440,"text":404},{"id":588,"depth":440,"text":589},{"id":773,"depth":440,"text":774},{"id":893,"depth":440,"text":894},{"id":915,"depth":440,"text":916},[86,947,948,949,950,951,952,953,954,955],"Artificial Intelligence","AMD Radeon","AI Inference","GPU Performance","Inference Latency","Inference Optimization","Large Language Models","vLLM","Cost Efficiency","2026-09-04T09:00:00","Paiton serves AMD’s Qwen3.8 27B at 39.77 output tokens\u002Fs on one Radeon AI PRO R9700, delivering 21% more throughput and 17.4% lower modeled cost.","md",null,"\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-r9700\u002F00-featured-paiton-r9700.webp",{},true,"https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","\u002Fblog\u002Fpaiton-qwen38-radeon-ai-pro-r9700",{"title":5,"description":957},"paiton-qwen38-radeon-ai-pro-r9700","blog\u002Fpaiton-qwen38-radeon-ai-pro-r9700","DH3nF8ztVdpbfdkqPQNUhH2eAYbOlr4j29LEMG4xl3Q",[970,981,992,1004,1014,1023,1025,1039,1070,1082,1104,1122,1141,1159,1177,1193,1205,1221,1236,1248,1257,1265,1280,1292,1303,1314,1324,1337,1347,1360,1371,1381,1392,1401,1413,1424,1433],{"path":971,"title":972,"description":973,"date":974,"slug":975,"image":976,"originalUrl":977,"categories":978},"\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700","Qwen3.8: 400.7 tok\u002Fs on R9700 | Paiton","Qwen3.8 on one R9700: 400.7 aggregate tok\u002Fs with ROCm 10 and vLLM 0.29, plus public 200K\u002F220K chat profiles. Benchmarks, limits and launch commands.","2026-09-16T07:30:00Z","paiton-qwen38-mxfp4-dflash2-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-qwen38-mxfp4\u002Fupdate-2026-09-19\u002Fhero.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-mxfp4-dflash2-r9700",[86,948,954,979,952,980],"Qwen3.8","DFlash2",{"path":982,"title":983,"description":984,"date":985,"slug":986,"image":987,"originalUrl":988,"categories":989},"\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700","Qwen3.8 GGUF in vLLM: Faster Responses on One Radeon","Run the original NEO CODER MAX GGUF in vLLM with Paiton on an R9700. Explore measured latency gains, image input and local deployment.","2026-09-14T07:30:00Z","paiton-qwen38-neo-gguf-vllm-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-neo-gguf\u002F00-hero-neo-gguf-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-qwen38-neo-gguf-vllm-r9700",[86,948,990,991,954],"Local AI","GGUF",{"path":993,"title":994,"description":995,"date":996,"slug":997,"image":998,"originalUrl":999,"categories":1000},"\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700","MiniMax H3 on Radeon: 15-Second Video With Native Sound","Paiton generates a 15-second MiniMax H3 video with stereo audio on one Radeon AI PRO R9700 in 5m 33s, with 16.7% lower latency than matched stock.","2026-09-09T07:30:00Z","paiton-minimax-h3-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-minimax-h3\u002F00-featured-minimax-h3-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-minimax-h3-radeon-ai-pro-r9700",[86,948,990,1001,1002,1003],"Video Generation","MiniMax H3","ComfyUI",{"path":1005,"title":1006,"description":1007,"date":1008,"slug":1009,"image":1010,"originalUrl":1005,"categories":1011},"\u002Fblog\u002Fpaiton-flux2-klein-radeon-ai-pro-r9700","Local FLUX.2 klein on Radeon AI PRO R9700: Faster Image Generation with Less VRAM","Paiton generates 1024×1024 FLUX.2 klein images in 1.054 seconds on an R9700: 16.2% lower warm latency and 33.4% lower peak Torch allocation.","2026-09-07T09:00:00","paiton-flux2-klein-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-flux2-klein\u002Ffox-paiton.webp",[86,948,990,1012,1013,1003],"Image Generation","FLUX",{"path":1015,"title":1016,"description":1017,"date":1018,"slug":1019,"image":1020,"originalUrl":1021,"categories":1022},"\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700","Ornith 1.5 at 44.6 tok\u002Fs on One Radeon AI PRO R9700","Paiton serves Ornith 1.5 35B A3B at 44.63 output tok\u002Fs on one Radeon AI PRO R9700, 27% faster than tuned stock vLLM, with 21.3% lower modeled cost.","2026-09-05T09:00:00","paiton-ornith15-radeon-ai-pro-r9700","\u002Fasset\u002Fimages\u002Fblog\u002Fpaiton-ornith15\u002F00-featured-ornith15-r9700.webp","https:\u002F\u002Feliovp.com\u002Fblog\u002Fpaiton-ornith15-radeon-ai-pro-r9700",[86,947,948,949,950,951,952,953,954,955],{"path":964,"title":5,"description":957,"date":956,"slug":966,"image":960,"originalUrl":963,"categories":1024},[86,947,948,949,950,951,952,953,954,955],{"path":1026,"title":1027,"description":1028,"date":1029,"slug":1030,"image":1031,"originalUrl":959,"categories":1032},"\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt","AI Data Center Power Requirements: The GPU\u002FMW Illusion","Why do AI data-center proposals quote different GPU capacities? See how PUE, peak loads, storage, networking and cooling determine deployable compute.","2026-07-27T23:52:00","ai-data-center-power-requirements-gpu-per-megawatt","\u002Fasset\u002Fimages\u002Fblog\u002Fai-data-center-power-requirements-gpu-per-megawatt\u002Fgpu-per-megawatt-illusion.webp",[1033,1034,1035,1036,1037,1038],"All","AI Infrastructure","Data Centers","ModFlex","HPC","AMD Helios",{"path":1040,"title":1041,"description":1042,"date":1043,"slug":1044,"image":1045,"originalUrl":1046,"categories":1047},"\u002Fblog\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","Wan2.2 Video Generation: Paiton on AMD MI355X","Compare Wan2.2-T2V-A14B video generation on AMD MI355X with Paiton and NVIDIA B200 using Diffusers, and explore our diffusion optimization approach.","2026-06-10T14:04:04","paiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonwan2.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-returns-to-its-diffusion-roots-optimizing-wan2-2-t2v-a14b-on-amd-mi355x\u002F",[1033,947,86,1048,1049,1050,1051,1052,1053,1054,1055,1056,793,1057,1058,1059,1060,1061,1062,1063,86,1064,1065,1066,1067,1068,1069],"14B","AMD","B200","Benchmarks","Blackwell","Compute","Diffusion","Eliovp","Generative AI","Hardware","Inference","Instinct","MI355x","NVidia","On-Premise","Optimization","Sovereign AI","T2V","Text-to-Video","Tuning","Video-Generation","Wan2.2",{"path":1071,"title":1072,"description":1073,"date":1074,"slug":1075,"image":1076,"originalUrl":1077,"categories":1078},"\u002Fblog\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","ElioVP in De Tijd: Chip Optimization and Data Centers","Read about De Tijd's coverage of ElioVP, from its origins in chip optimization to its work on modular data centers and high-density cooling.","2026-02-10T20:48:12","from-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fphysicalnewspaper.webp","https:\u002F\u002Feliovp.com\u002Ffrom-the-attic-to-the-front-page-eliovp-recognized-as-a-pioneer-in-chip-optimization-data-center-infrastructure\u002F",[1033,947,1079,1080,1049,1081,1079,1061],"Modular DC","Uncategorized","De Tijd",{"path":1083,"title":1084,"description":1085,"date":1086,"slug":1087,"image":1088,"originalUrl":1089,"categories":1090},"\u002Fblog\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","AI Privacy: A Strategic Priority for Benelux Businesses","Explore generative AI privacy risks, trust, data retention and governance, and why Benelux businesses need a strategic approach to secure AI.","2026-01-29T13:51:11","privacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fheaderimage.webp","https:\u002F\u002Feliovp.com\u002Fprivacy-is-geen-it-probleem-meer-het-is-een-strategische-prioriteit\u002F",[1033,947,1091,1080,1092,1093,1094,1095,1096,1097,1098,1099,1055,1100,1056,1101,1102,1103],"Trending","AI Act","Anthropomorphism","AVG","Benelux","ChatGPT","Cybersecurity","Data Governance","Data Security","GDPR","Microsoft Copilot","Privacy","Shadow AI",{"path":1105,"title":1106,"description":1107,"date":1108,"slug":1109,"image":1110,"originalUrl":1111,"categories":1112},"\u002Fblog\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","Why We Do Not Use itsme: Privacy and Data Sovereignty","Why ElioVP does not use itsme: our assessment of identity metadata, cloud dependence, data sovereignty and authentication risks.","2025-11-27T09:32:14","itsme-bij-ons-is-het-its-not-me-en-dit-is-waarom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontimage.webp","https:\u002F\u002Feliovp.com\u002Fitsme-bij-ons-is-het-its-not-me-en-dit-is-waarom\u002F",[1033,1113,1114,1115,1097,1116,1117,1118,1100,1119,1120,1121,1102],"AWS","Belgian Mobile ID","Cloud Act","Data Sovereignty","Digital Identity","eIDAS","itsme","Liberty Global","MyGov.be",{"path":1123,"title":1124,"description":1125,"date":1126,"slug":1127,"image":1128,"originalUrl":1129,"categories":1130},"\u002Fblog\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025","Field Report. The Reality of Building Agentic AI in 2025","Lessons from building on-premise AI agents in 2025 cover workflow design, observability, model training, hallucinations and GPU memory limits.","2025-11-25T14:03:39","field-report-the-reality-of-building-agentic-ai-in-2025","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffieldreport.webp","https:\u002F\u002Feliovp.com\u002Ffield-report-the-reality-of-building-agentic-ai-in-2025\u002F",[1033,947,1131,1091,1132,1133,1134,1135,1136,1137,1138,1139,1064,1140],"Solutions","Agentic AI","AI Engineering","AI Strategy","Autonomous Agents","Enterprise AI","Local LLM","Model Fine-Tuning","On-Premise AI","VRAM Optimization",{"path":1142,"title":1143,"description":1144,"date":1145,"slug":1146,"image":1147,"originalUrl":1148,"categories":1149},"\u002Fblog\u002Fthe-synthetic-unicorn-bubble","The Synthetic Unicorn Bubble","An analysis of AI neocloud investment risks, examining circular financing, infrastructure claims, contract terms and due diligence.","2025-11-24T19:22:28","the-synthetic-unicorn-bubble","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsyntheticunicorn.webp","https:\u002F\u002Feliovp.com\u002Fthe-synthetic-unicorn-bubble\u002F",[1033,947,1091,1034,1150,1151,1152,1153,1154,1155,1156,1157,1158],"AI Neocloud","Circular Financing","GPU Cloud","Investment Risks","Startup Valuation","Synthetic Bubble","Tech Analysis","Vaporware","Venture Capital",{"path":1160,"title":1161,"description":1162,"date":1163,"slug":1164,"image":1165,"originalUrl":1166,"categories":1167},"\u002Fblog\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","NVIDIA GB300 NVL72: A Four-Month Modular Data Center Plan","Explore a modular data center design for NVIDIA GB300 NVL72, covering redundant power, hybrid cooling and a four-month deployment plan.","2025-11-20T14:10:19","building-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fsuperpodmodflexfrontimage.webp","https:\u002F\u002Feliovp.com\u002Fbuilding-the-engine-for-the-ai-race-the-4-month-path-to-nvidia-gb300-nvl72-power\u002F",[1033,1079,1080,1168,1034,1169,1170,1171,1172,1173,1174,1175,1176],"150kW Rack","DLC","High Density","Liquid Cooling","Modular Data Center","NVIDIA Blackwell Ultra","NVIDIA GB300","NVL72","Rapid Deployment",{"path":1178,"title":1179,"description":1180,"date":1181,"slug":1182,"image":1183,"originalUrl":1184,"categories":1185},"\u002Fblog\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential","Why “CUDA” Translation Won’t Unlock AMD’s Real Potential","Why CUDA compatibility is not the same as AMD performance: explore ROCm, HIP, kernel tuning and the case for hardware-specific optimization.","2025-11-12T14:48:37","why-cuda-translation-wont-unlock-amds-real-potential","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fchatgpt-image-nov-11-2025-09_16_10-pm-1.webp","https:\u002F\u002Feliovp.com\u002Fwhy-cuda-translation-wont-unlock-amds-real-potential\u002F",[1033,947,86,1080,1186,947,1187,1188,1189,1190,1191,1192,86,822],"AMD MI300X","CUDA Translation","FP8","GPU Optimization","High Performance Computing","HIP","Kernel Tuning",{"path":1194,"title":1195,"description":1196,"date":1197,"slug":1198,"image":1199,"originalUrl":1200,"categories":1201},"\u002Fblog\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference","Paiton: The Simplest Way to Supercharge AI Inference","Learn how Paiton integrates with existing inference stacks, with AMD MI300X benchmark results and performance-per-dollar comparisons.","2025-11-11T10:31:22","paiton-the-simplest-way-to-supercharge-ai-inference","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton-powaaah.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-the-simplest-way-to-supercharge-ai-inference\u002F",[1033,947,86,949,1202,1186,955,1203,952,1192,86,1204,954],"AMD Instinct","High Throughput","SGLang",{"path":1206,"title":1207,"description":1208,"date":1209,"slug":1210,"image":1211,"originalUrl":1212,"categories":1213},"\u002Fblog\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","Paiton MoE Benchmarks: MI300X vs H200 and B200","Compare Qwen3-30B-A3B MoE inference with Paiton on MI300X against H200 and B200, including throughput and cost per million tokens.","2025-09-26T13:36:18","stop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhulkvshulkpaitonwins.webp","https:\u002F\u002Feliovp.com\u002Fstop-overpaying-paiton-mi300x-moe-beats-h200-b200-on-1m-tokens\u002F",[1033,947,86,1214,1186,1215,952,1216,1217,1218,1219,86,1220],"AI Benchmarks","Cost per Token","Mixture of Experts","MoE","NVIDIA B200","NVIDIA H200","Qwen3",{"path":1222,"title":1223,"description":1224,"date":1225,"slug":1226,"image":1227,"originalUrl":1228,"categories":1229},"\u002Fblog\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","Local Agentic AI: From Inbox to Action","Local-first AI agents turn email, documents and images into tickets, reports and actions, using models tailored to your data and systems.","2025-09-16T13:09:00","agentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ffrontfotoblog.webp","https:\u002F\u002Feliovp.com\u002Fagentic-ai-but-make-it-local-from-inbox-to-insight-to-action-en\u002F",[1033,947,1131,1080,1132,1230,1231,1232,1233,1137,1139,1064,1234,1235],"Damage Detection","Document Processing","Email Automation","Invoice Extraction","Ticket Automation","Workflow Automation",{"path":1237,"title":1238,"description":1239,"date":1240,"slug":1241,"image":1242,"originalUrl":1243,"categories":1244},"\u002Fblog\u002Fmi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","MI300X FP8 Benchmarks: GPU Partitioning with Paiton","Explore Llama 3.1 8B FP8 benchmarks on partitioned MI300X GPUs with Paiton, comparing throughput and latency against NVIDIA H200 and B200.","2025-07-31T13:32:57","mi300x-fp8-data-parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fimage-2-1.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-fp8-data%e2%80%91parallel-benchmarks-8-64-gpus-h200-left-behind-b200-within-reach\u002F",[1033,947,86,1080,1245,1049,1050,1246,1247,1060,1061,86,954],"AI","H200","MI300X",{"path":1249,"title":1250,"description":1251,"date":1252,"slug":1253,"image":1254,"originalUrl":1255,"categories":1256},"\u002Fblog\u002Fapplicable-ai-for-businesses","Applicable AI for Businesses","Explore ElioVP's approach to local AI for business workflows, including custom model training and automated damage detection for logistics.","2025-07-09T21:35:30","applicable-ai-for-businesses","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fscherm_afbeelding-2025-07-09-om-23.30.17.webp","https:\u002F\u002Feliovp.com\u002Fapplicable-ai-for-businesses\u002F",[1033,947,1131,1132,1230,1231,1232,1233,1137,1139,1064,1234,1235],{"path":1258,"title":1259,"description":1260,"date":1261,"slug":1262,"image":414,"originalUrl":1263,"categories":1264},"\u002Fblog\u002Fintroducing-paitons-free-evaluation-models","Introducing Paiton’s Free Evaluation Models","Test Paiton with free evaluation models for AMD GPUs. Compare text, vision and image generation performance using your own workloads.","2025-07-07T11:26:13","introducing-paitons-free-evaluation-models","https:\u002F\u002Feliovp.com\u002Fintroducing-paitons-free-evaluation-models\u002F",[1033,947,86],{"path":1266,"title":1267,"description":1268,"date":1269,"slug":1270,"image":1271,"originalUrl":1272,"categories":1273},"\u002Fblog\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","Llama 3.1 405B: Faster Startup with Paiton on MI300X","See Paiton benchmarks for Llama 3.1 405B on eight AMD MI300X GPUs, covering model startup, tensor parallelism, throughput and latency.","2025-06-12T20:15:23","paiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fservingscreenshot.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-dramatically-faster-startup-and-performance-for-llama-3-1-405b\u002F",[1033,947,86,1080,949,1186,1274,1188,1275,1276,1277,86,1278,1279],"Cold Start","Graph Compilation","Llama 3.1 405B","LLM Optimization","Startup Latency","Tensor Parallelism",{"path":1281,"title":1282,"description":1283,"date":1284,"slug":1285,"image":1286,"originalUrl":1287,"categories":1288},"\u002Fblog\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x","Paiton FP8 Beats NVIDIA’s H200 on AMD’s MI300X","Compare Paiton on AMD MI300X with NVIDIA H200 for Llama 3.1 70B FP8, including throughput, first-token delay and latency across batch sizes.","2025-06-08T19:12:40","paiton-fp8-beats-nvidias-h200-on-amds-mi300x","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fblognewfp8.webp","https:\u002F\u002Feliovp.com\u002Fpaiton-fp8-beats-nvidias-h200-on-amds-mi300x\u002F",[1033,947,86,1080,1186,1289,1136,1056,950,951,953,1276,1290,1291],"Cold Start Optimization","Model Serving","vLLM Optimization",{"path":1293,"title":1294,"description":1295,"date":1296,"slug":1297,"image":1298,"originalUrl":1299,"categories":1300},"\u002Fblog\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","MI300X vs H200 vs RX 7900 XTX vs Tenstorrent n300s with vLLM","Compare MI300X, H200, RX 7900 XTX and Tenstorrent n300s on Llama 3 8B with vLLM, including throughput, modeled token costs and hardware limits.","2025-05-09T14:03:58","mi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisontenstor.webp","https:\u002F\u002Feliovp.com\u002Fmi300x-vs-h200-vs-rx-7900-xtx-vs-tenstorrent-n300s-with-vllm\u002F",[1033,947,86,1131,1080,1049,1247,1061,1301,1302],"RX7900XTX","tenstorrent",{"path":1304,"title":1305,"description":1306,"date":1307,"slug":1308,"image":1309,"originalUrl":1310,"categories":1311},"\u002Fblog\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","ClusterP&L: Financial Modeling for GPU Clusters","Explore how ClusterP&L models GPU cluster costs, profitability and investment scenarios, with ROI metrics, risk simulations and exportable reports.","2025-05-03T10:52:22","clusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fcomparisonscenarios.webp","https:\u002F\u002Feliovp.com\u002Fclusterpl-empowering-gpu-cluster-investors-with-real-world-financial-insights\u002F",[1033,947,1079,1131,1050,1246,1312,1061,1313],"MI325x","pnl calculator",{"path":1315,"title":1316,"description":1317,"date":1318,"slug":1319,"image":1320,"originalUrl":1321,"categories":1322},"\u002Fblog\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","AMD MI300X vs. NVIDIA H200: Qwen3-32B with Paiton","Compare Qwen3-32B benchmarks on Paiton-optimized AMD MI300X and NVIDIA H200, covering throughput, latency and hardware costs.","2025-05-02T21:10:30","cranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002F3ac59a73-2466-4422-b7e5-ef2e4a8ca58e.webp","https:\u002F\u002Feliovp.com\u002Fcranking-out-faster-tokens-for-fewer-dollars-amd-mi300x-vs-nvidia-h200\u002F",[1033,947,86,1245,1049,1246,1323,1061,86,954],"MI300",{"path":1325,"title":1326,"description":1327,"date":1328,"slug":1329,"image":1330,"originalUrl":1331,"categories":1332},"\u002Fblog\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","Modular Data Centers for NVIDIA NVL: 1 to 2 MW","Explore modular data center designs for NVIDIA NVL systems, covering power capacity, liquid cooling, redundancy and deployment planning.","2025-05-02T14:09:59","power-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Feliovp_critical-1mw-pod_rev-2_transparent.webp","https:\u002F\u002Feliovp.com\u002Fpower-meets-precision-high-density-modular-data-center-for-nvidia-nvl-deployments-1-2-mw\u002F",[1033,1079,1333,1034,1170,1037,1171,1172,1334,1335,1175,1336],"1-2MW Data Center","NVIDIA Blackwell","NVIDIA NVL","Precision Cooling",{"path":1338,"title":1339,"description":1340,"date":1341,"slug":1342,"image":1343,"originalUrl":1344,"categories":1345},"\u002Fblog\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","Examining AI agents in the medical field: AI that speaks DICOM","Explore a local AI agent demo for DICOM workflows, from patient and study retrieval to a comparison of vision models using anonymized medical images.","2025-04-11T14:45:48","examining-ai-agents-in-the-medical-field-ai-that-speaks-dicom","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fhealthcareblog-1.webp","https:\u002F\u002Feliovp.com\u002Fexamining-ai-agents-in-the-medical-field-ai-that-speaks-dicom\u002F",[1033,947,1131,1080,1245,1049,1346],"Healthcare",{"path":1348,"title":1349,"description":1350,"date":1351,"slug":1352,"image":1353,"originalUrl":1354,"categories":1355},"\u002Fblog\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","U.S. Tariffs and AI Supply Chain Resilience: April 2025","Read ElioVP's April 2025 perspective on U.S. tariffs and supply chain resilience for AI servers, HPC systems and modular data centers.","2025-04-04T10:01:27","eliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftariffsshipping.webp","https:\u002F\u002Feliovp.com\u002Feliovp-bv-your-trusted-partner-for-supply-chain-resilience-amidst-new-u-s-tariffs\u002F",[1033,1091,1245,1049,1356,1357,1358,1359],"import","Taiwan","Tariffs","Trump",{"path":1361,"title":1362,"description":1363,"date":1364,"slug":1365,"image":1366,"originalUrl":1367,"categories":1368},"\u002Fblog\u002Fwhy-ai-agents-are-the-future","Why AI Agents Are the Future","Explore AI agents for ERP, CRM, finance and customer support, with practical use cases and a path from workflow assessment to pilot and deployment.","2025-03-23T22:06:59","why-ai-agents-are-the-future","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ferp2.jpeg","https:\u002F\u002Feliovp.com\u002Fwhy-ai-agents-are-the-future\u002F",[1033,947,1131,1245,1369,1370],"AI Agents","ERP",{"path":1372,"title":1373,"description":1374,"date":1375,"slug":1376,"image":1377,"originalUrl":1378,"categories":1379},"\u002Fblog\u002Fthe-rise-of-open-source-ai-model-optimization","The Rise of Open-Source AI Model Optimization","Explore open-source AI optimization trends, from quantization and mixture-of-experts models to hardware-aware tuning, RAG and edge deployment.","2025-03-22T20:59:23","the-rise-of-open-source-ai-model-optimization","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Friseofopensource.jpeg","https:\u002F\u002Feliovp.com\u002Fthe-rise-of-open-source-ai-model-optimization\u002F",[1033,947,1091,1380,1049,793,1061],"AI news",{"path":1382,"title":1383,"description":1384,"date":1385,"slug":1386,"image":1387,"originalUrl":1388,"categories":1389},"\u002Fblog\u002Fintroducing-our-benchmarking-tool-powered-by-dstack","Introducing Our Benchmarking Tool: Powered by dstack","Explore our dstack-powered tool for reproducible vLLM benchmarks, automated parameter sweeps and performance reports across local and cloud GPUs.","2025-03-20T14:21:59","introducing-our-benchmarking-tool-powered-by-dstack","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fbenchmarktool.jpeg","https:\u002F\u002Feliovp.com\u002Fintroducing-our-benchmarking-tool-powered-by-dstack\u002F",[1033,947,86,1245,1049,1390,1391,1247,86],"benchmark","LLM",{"path":1393,"title":1394,"description":1395,"date":1396,"slug":1397,"image":1398,"originalUrl":1399,"categories":1400},"\u002Fblog\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","Optimizing QwQ-32B (by Qwen): AMD MI300X vs. NVIDIA H200","Compare QwQ-32B throughput and latency on AMD MI300X with Paiton and NVIDIA H200, from small batches to higher concurrency.","2025-03-19T21:41:44","optimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaiton4.jpeg","https:\u002F\u002Feliovp.com\u002Foptimizing-qwq-32b-by-qwen-amd-mi300x-vs-nvidia-h200\u002F",[1033,947,86],{"path":1402,"title":1403,"description":1404,"date":1405,"slug":1406,"image":1407,"originalUrl":1408,"categories":1409},"\u002Fblog\u002Feliovp-featured-on-amd-tech-talk-podcast","Eliovp Featured on AMD “Tech Talk” Podcast","Listen to Elio Van Puyvelde and Jim Greene on AMD's Tech Talk podcast, discussing ElioVP's origins and its AI hardware and software services.","2025-03-19T07:53:39","eliovp-featured-on-amd-tech-talk-podcast","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Ftechtalkjimgreene.jpeg","https:\u002F\u002Feliovp.com\u002Feliovp-featured-on-amd-tech-talk-podcast\u002F",[1033,1049,1410,1411,1412],"Jim Greene","Podcast","Tech Talk",{"path":1414,"title":1415,"description":1416,"date":1417,"slug":1418,"image":1419,"originalUrl":1420,"categories":1421},"\u002Fblog\u002Ffurther-optimizing-amd-powered-inference-with-paiton","Further Optimizing AMD-Powered Inference with Paiton","Explore Paiton's DeepSeek R1 Distill Llama 8B benchmarks on AMD MI300X, focusing on throughput and first-token latency at smaller batch sizes.","2025-03-13T06:18:30","further-optimizing-amd-powered-inference-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost3.webp","https:\u002F\u002Feliovp.com\u002Ffurther-optimizing-amd-powered-inference-with-paiton\u002F",[1033,947,86,1049,1422,1423,1246,1247,1312,86,954],"Deepseek","H100",{"path":1425,"title":1426,"description":1427,"date":1428,"slug":1429,"image":1430,"originalUrl":1431,"categories":1432},"\u002Fblog\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","Paiton Benchmarks: DeepSeek R1 Distill Llama 3.1 8B","Compare stock and Paiton-optimized DeepSeek R1 Distill Llama 3.1 8B on AMD MI300X, with throughput and latency benchmarks across batch sizes.","2025-01-31T09:11:02","a-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost2.webp","https:\u002F\u002Feliovp.com\u002Fa-first-look-at-paiton-in-action-deepseek-r1-distill-llama-3-1-8b\u002F",[1033,947,86,1049,1422,1423,1246,1247,1312,86,954],{"path":1434,"title":1435,"description":1436,"date":1437,"slug":1438,"image":1439,"originalUrl":1440,"categories":1441},"\u002Fblog\u002Fai-model-optimization-with-paiton","AI Model Optimization with Paiton","Learn how Paiton uses model compilation, custom kernels and kernel fusion to optimize AI inference on AMD GPUs.","2025-01-30T19:53:25","ai-model-optimization-with-paiton","\u002Fasset\u002Fimages\u002Fblog\u002Fimported\u002Fpaitonpost1.webp","https:\u002F\u002Feliovp.com\u002Fai-model-optimization-with-paiton\u002F",[1033,947,86,1049,1423,1246,1247,1312,86,954],1789853166465]