Ten sam model (Qwen3.8-27B), ta sama metoda: prompt 2 849 tokenów, 600 tokenów odpowiedzi, rozumowanie włączone, średnie z powtórzeń. Jaśniejszy pasek — stan przed zmianą, pełny — po zmianie. Pełne tabele są w benchmarku, a co dokładnie zmienialiśmy — w punktach niżej.
--enforce-eager (grafy CUDA), spekulacja MTP i większa kolejka podniosły generowanie z 8,4 do 28,9 tok/s — ×3,4 bez zmiany sprzętu. Kolejne +45 % dał limit mocy 220 W zamiast 180 W na maszynie, której zasilacz go wytrzymuje (Z440: 32,3 → 46,8 tok/s).llm-scaler): grafy XPU weszły w lipcu 2026 jako eksperymentalne i w przykładach dalej są wyłączone, MTP dla Qwen 27B doszło 4 sierpnia, poprawki cache prefiksu i współbieżności — 9 września, a MTP przy wielu użytkownikach wciąż się sypie. Nasza B70 pracowała od 2 września; przez pierwsze trzy tygodnie jedyną stabilną drogą było llama.cpp. Kupując Intela, kupujesz tańszą pamięć od razu, a wydajność — z opóźnieniem, razem z oprogramowaniem.Mierzone wyłącznie GPU i pakiet CPU, z liczników energii sterownika i RAPL. Nie obejmuje płyty, dysków, wentylatorów obudowy, strat sekcji zasilania ani sprawności zasilacza — rzeczywisty pobór z gniazdka jest wyższy.
Ta sama maszyna, ta sama karta, ten sam limit mocy 180 W i ta sama blokada zegara 1500 MHz. Zmieniona wyłącznie konfiguracja silnika vLLM. Pomiar na Z420 z 24.09.2026, średnie z dwóch powtórzeń każdego przebiegu.
Ta sama maszyna, ta sama konfiguracja vLLM, zmieniany wyłącznie limit mocy karty. Blokada zegara na 1500 MHz pozostaje włączona przez cały pomiar.
--max-num-seqs 4 i
--gpu-memory-utilization 0.94. Pula cache KV wynosi wtedy 111 287 tokenów,
czyli 1,7 zapytania pełnej długości naraz albo odpowiednio więcej krótszych.--max-num-seqs 8), bo przy tych ustawieniach silnik
obsługuje więcej zapytań naraz. Prędkość dla jednego użytkownika jest w obu wariantach taka sama —
28,9 tok/s przy 32 tys. i 28,9 tok/s przy 64 tys. — więc wiersze jednoosobowe obowiązują dla obu.
Kluczowa dla pamięci jest wielkość kolejki, nie samo okno: grafy CUDA zajmują tym więcej, im więcej
rozmiarów wsadu muszą przechwycić.llm-scaler-vllm:0.26.0-b2, GPTQ int4, cache KV fp8, grafy XPU, cache prefiksu), zmierzony tym samym promptem (2 849 tok.) i tą samą metodą co Z820 i Z440, przy limicie 250 W. Wcześniejsze liczby z llama.cpp — w sekcji o B70 niżej, jako historia. Pobór w spoczynku pochodzi z 10.09.--enforce-eager (czyli włączone grafy
CUDA), zachowana spekulacja MTP, kolejka --max-num-seqs podniesiona z 3 do 8, włączony
--enable-chunked-prefill, okno kontekstu skrócone do 32 tys. tokenów, żeby starczyło
pamięci na grafy. Żadnej zmiany w sprzęcie.--enforce-eager host musi ręcznie wystrzelić każde jądro obliczeniowe.
Po włączeniu grafów karta pracuje na limicie i to ona jest wąskim gardłem, a nie maszyna.
Stąd też trzykrotny skok wydajności bez wymiany czegokolwiek.Do 4 października 2026 czwartą kolumną benchmarku była pojedyncza RTX 3090 w HP Z440 (Xeon E5-2683 v3, 62 GB DDR4). Zastąpiły ją cztery 3090 z Flash-Next; wyniki zostają tutaj, bo to wciąż najtańsza droga do lokalnego modelu. Qwen3.8-27B int4 AutoRound, KV fp8, okno 65 536, max-num-seqs 8, vLLM 0.20.1, pomiar 24.09.2026.
| 3090 · Z440 · 180 W | |
|---|---|
| 1 użytkownik — czas / tok/s | 22,7 s / 26,4 |
| Samo generowanie | 32,3 tok/s |
| Czas do 1. tokenu / czytanie promptu | 4,2 s / 675 tok/s |
| 3 użytkowników — łącznie / na osobę | 48,5 / 16,2 tok/s |
| 8 użytkowników — łącznie | 45,8 tok/s |
| Pobór GPU / CPU / razem | 179 W / 43 W / ~222 W |
| Prąd pod obciążeniem — zmierzone / z gniazdka | 160 zł / 239 zł mies. |
| Sprzęt na 1 tok/s (1 / 3 użytk.) | 171 zł / 93 zł |
| Wydajność na wat (1 / 3 użytk.) | 0,147 / 0,271 tok/s/W |
| Limit mocy | Generowanie 1 użytkownik | 1 użytkownik tok/s | 3 naraz łącznie | 8 naraz łącznie | Czas do 1. tokenu | Czytanie promptu | Temp. |
|---|---|---|---|---|---|---|---|
| 180 W | 32,3 tok/s | 26,4 | 48,5 tok/s | 45,8 tok/s | 4,2 s | 675 tok/s | 74 °C |
| 200 W | 38,5 tok/s | 31,3 | 60,1 tok/s | 57,7 tok/s | 3,6 s | 790 tok/s | 79 °C |
| 220 W | 46,8 tok/s | 37,7 | 71,9 tok/s | 69,1 tok/s | 3,1 s | 919 tok/s | 82 °C |
Domyślna konfiguracja obu silników ogranicza liczbę równolegle przetwarzanych sekwencji do jednej. To nie jest ograniczenie sprzętu — to jedna wartość w konfiguracji, która nie kosztuje dodatkowego VRAM, bo pula KV jest przydzielana raz, przy starcie.
| Silnik | Parametr | Zmiana | Przepustowość, 3 użytkowników |
|---|---|---|---|
| vLLM | --max-num-seqs | 3 → 24 | 33,3 → 63,7 tok/s przy 8 użytkownikach (×1,91) |
| llama.cpp | --parallel | 1 → 3 | 18,95 → 27,05 tok/s (×1,43) |
Kolejka działa jednak tylko wtedy, gdy nie ma spekulacji. Z włączoną spekulacją MTP podniesienie kolejki z 3 do 8 nie zmieniło prawie nic (45,6 → 46,1 tok/s przy ośmiu użytkownikach) — spekulacja zużywa ten sam zapas obliczeniowy, z którego żyje batching.
Dlaczego Arc skaluje słabiej (×1,44 wobec ×1,91): przy jednym strumieniu Arc już wykorzystuje moc obliczeniową — przepustowość na użytkownika spada z 19 do 10 tok/s przy trzech. RTX 3090 przy jednym strumieniu jest ograniczona przepustowością pamięci i ma zapas obliczeniowy. Paradoksalnie: im szybsza karta na pojedynczym strumieniu, tym mniej zyskuje na batchingu.
Z włączonym rozumowaniem model zużywa około 350 tokenów na sam blok myślenia. Przy limicie 600 tokenów odpowiedź wraca losowo pusta — w jednym przebiegu 996 znaków treści, w kolejnym zero, przy identycznej konfiguracji. To wariancja na granicy budżetu, nie usterka.
Przy 2500 tokenach: 1407 znaków rozumowania i 6986 znaków właściwej odpowiedzi. Praktyczne minimum to 1500–2000 tokenów.
Karta 3090 pracuje z limitem mocy 180 W. Dołożenie blokady zegara na 1500 MHz — mimo że karta potrafi 1980 MHz — podniosło przepustowość z 6,8 do 8,0 tok/s przy jednym użytkowniku.
Przy samym limicie mocy karta boostuje do maksimum, natychmiast uderza w sufit poboru i zostaje ostro zdławiona — pracuje w piłę. Przypięta do niższego, stałego zegara mieści się w budżecie mocy w sposób ciągły. Stabilny niższy zegar bije oscylujący wyższy, a przy okazji znosi mikrosekundowe szpilki poboru, na które sam limit mocy nie działa.
Na Arcu kontekst 262 144 tokenów przy KV w q8_0 zajmuje około 11,5 GiB ponad same wagi modelu (14,3 GiB). Podniesienie całkowitego kontekstu do 393 216 i rozdzielenie go na trzy sloty daje po 131 072 tokeny na użytkownika:
przed: 25,76 GiB / 32 1 slot × 262 144 po: 30,65 GiB / 32 3 sloty × 131 072 bez błędów alokacji
Na vLLM zależność działa w drugą stronę: włączenie grafów CUDA ścięło
dostępny KV cache z 4,61 do 2,24 GiB, a na 64 tys. tokenów kontekstu potrzeba 2,3 GiB —
silnik odmówił startu. Przy włączonej spekulacji grafy schodzą do trybu częściowego
(PIECEWISE), bo backend FlashInfer nie obsługuje pełnego.
Mimo trybu częściowego zysk jest duży: 13,3 → 31,7 tok/s na użytkownika
na tej samej maszynie i przy tym samym limicie mocy. Sama spekulacja MTP odpowiada za ×1,7 —
bez niej te same ustawienia dają 19,0 tok/s. Pułapka jest w rachunku pamięci: grafy alokują
się poza budżetem --gpu-memory-utilization, więc trzeba obniżyć
budżet albo skrócić kontekst. Przy kolejce 24 i włączonej spekulacji nie udało się znaleźć
ustawienia, przy którym silnik w ogóle wstaje na 24 GB.
Sterownik xe wystawia w systemie plików to, czego karty GeForce nie
pokazują: liczniki energii osobno dla karty i pakietu, temperatury rdzenia, pamięci,
kontrolera i magistrali PCIe, a do tego szesnaście kanałów VRAM osobno,
obroty wentylatora oraz limit i próg krytyczny mocy.
RTX 3090 nie raportuje żadnych napięć — to celowe ograniczenie kart konsumenckich — ani temperatur poza rdzeniem. Jedynym zamiennikiem jest licznik hamowania mocy, rosnący, gdy karcie brakuje zasilania.
--enforce-eager potrafi zabrać ×3Bez grafów CUDA host odpala każde jądro osobno. Na starym Xeonie karta czekała na procesor i ciągnęła 131–138 W przy limicie 180 W. Po zdjęciu flagi: 8,4 → 28,9 tok/s na tej samej maszynie.
Jak rozpoznać: karta nie dochodzi do własnego limitu mocy pod obciążeniem. Wtedy wąskim gardłem jest host, nie GPU.
--gpu-memory-utilizationWłączenie grafów ścięło dostępny KV cache z 4,61 do 2,24 GiB, a silnik odmówił startu
przy 64k kontekstu. Im większe --max-num-seqs, tym więcej rozmiarów wsadu
do przechwycenia i tym więcej pamięci.
# działa na 1× 3090, 64k, pula KV 111 287 tok. --max-num-seqs 4 --gpu-memory-utilization 0.94 --max-model-len 65536
Przy kolejce 24 z włączoną spekulacją nie znaleźliśmy ustawienia, które w ogóle wstaje na 24 GB.
MTP daje ×1,7 dla jednego użytkownika. Ale z MTP podniesienie kolejki z 3 do 8 nic nie dało (45,6 → 46,1 tok/s przy 8 użytkownikach). Wybierz pod swój ruch: pojedynczy strumień → MTP, wiele osób naraz → kolejka bez spekulacji.
Z MTP grafy schodzą do trybu PIECEWISE (FlashInfer nie obsługuje pełnego).
I tak się opłaca.
Cache prefiksu sprawia, że drugi przebieg z tym samym promptem jest „magicznie” szybszy. Wstaw do każdego zapytania unikalny znacznik. Mierz strumieniowo, żeby oddzielić czas do pierwszego tokenu od generowania.
max_tokens = losowo pusta odpowiedźBlok myślenia to ~350 tokenów. Przy limicie 600 odpowiedź raz ma 996 znaków, raz zero, w identycznej konfiguracji. Minimum praktyczne: 1500–2000 tokenów.
--gpu-memory-utilization 0.96 startuje, a potem padaSilnik wstaje normalnie, a pod obciążeniem: torch.OutOfMemoryError: Tried to allocate 56.00 MiB. Najpodstępniejszy błąd, bo start wygląda na sukces. Nie idź powyżej 0.95. Podnoszenie budżetu dodatkowo zmniejsza miejsce na grafy — właściwą dźwignią jest --max-num-seqs.
--enable-chunked-prefill w vLLM 0.20.1 jest domyślnie włączony — w logu enable_chunked_prefill=True bez tej flagi--kv-cache-dtype przyjmuje tylko auto i warianty fp8. 4-bit KV ma llama.cppW logu startowym, w linii non-default args, stoi 'api_key': [...]. Kto czyta dziennik, ma klucz. Przed udostępnieniem logów albo maszyny — wymień klucz.
Ta sama 3090 na DDR3 i na DDR4: przy --enforce-eager różnica 35 %, po włączeniu grafów 12 %. Nie wymieniaj płyty ani pamięci, zanim nie sprawdzisz konfiguracji silnika.
Przy samym limicie 180 W karta boostuje, uderza w sufit i jest dławiona. Blokada na 1500 MHz: 6,8 → 8,0 tok/s, mimo niższego zegara. Znika też część szpilek poboru, na które limit mocy nie reaguje.
nvidia-smi -pl 180 nvidia-smi -lgc 1500,1500
180 → 220 W dało +45 % generowania, ale dopiero po włączeniu grafów. Przy eager zmiana limitu nic nie robiła, bo karta i tak nie dochodziła do 180 W.
Na 250 W w starej obudowie HP wyzwoliło się OCP zasilacza przy jednoczesnym obciążeniu CPU i GPU. Maszyna po prostu gaśnie, bez logów. Sam limit mocy ogranicza średnią w oknie milisekund, a nie mikrosekundowe szpilki przy batchingu — to one wywalają OCP. Na innej takiej maszynie padało już przy 200 W. Limit mocy to cecha zasilacza obudowy, nie karty — podnoś po 20 W.
--speculative-config '{"method": "mtp", "num_speculative_tokens": 2}' na Qwen3.8-27B w BF16, tensor-parallel 4: generowanie jednego użytkownika 40,3 → 55,5 tok/s. Przy trzech i ośmiu osobach zysk topnieje do 4–6 % (79,6 → 82,4 i 140,0 → 148,9 tok/s łącznie), bo karty i tak liczą pełnym wsadem. Akceptacja szkicu w godzinnym teście: 73 % — 82 % na pierwszej pozycji, 64 % na drugiej.
--disable-custom-all-reduce na czterech kartach PCIe nic nie zmieniavLLM przy starcie wypisuje: Custom allreduce is disabled because it's not supported on more than two PCIe-only GPUs. Bez NVLinka, przy więcej niż dwóch kartach, custom all-reduce jest wyłączony zawsze. Porównaliśmy „z flagą” i „bez flagi” i przy ośmiu użytkownikach wyszło 136 wobec 162 tok/s — to był jeden słaby przebieg na cztery, nie efekt flagi. Zanim przypiszesz różnicę zmianie, sprawdź w logu silnika, czy zmiana w ogóle zadziałała.
Ośmiu użytkowników przez 60 minut, każdy co rundę wkleja długi dokument, rozmowy do 50 tys. tokenów: 215 rund, zero błędów, zero restartów, zero Xid/AER, karty na limicie 220 W, najgorętsza 81 °C. Ale cache prefiksu trafiał tylko w 40 % — ośmiu rozmówców po ~50 tys. tokenów to prawie cała pula KV (431 tys. w BF16), więc bloki jednej rozmowy są wypychane przez zapytania innych, zanim ten sam użytkownik wróci. vLLM czyta wtedy całą rozmowę od nowa (~2 200 tok/s na czterech kartach) i czas do pierwszego tokenu rośnie do ~33 s. Dźwignią jest KV w fp8 — dwa razy większa pula kosztem jakości.
Ta sama maszyna Z820, nowy model. Flash-Next ma 125 mld parametrów głównych, 51 mld w tablicy n-gramów i głowicę MTP. Warstwy ekspertów mają po 512 ekspertów, z których na token pracuje 10 plus jeden wspólny — razem ok. 6 mld aktywnych parametrów. Trzy warstwy na cztery to Gated DeltaNet z pamięcią o stałym rozmiarze, co czwarta to Qwen Sparse Attention, która czyta tylko wybrane bloki tekstu. Natywnie 262 144 tokeny.
Ampere (sm86) nie ma sprzętowego FP8. Oficjalny vLLM obsługuje ten model od wersji 0.30.0 (22.09.2026), ale na 3090 tylko z pamięcią rozmów (KV) w 16 bitach — wtedy pula wychodzi ok. 400 tys. tokenów. Drogą jest kwantyzacja INT4 W4A16 liczona jądrami Marlin plus łatki społeczności, które dodają odczyt KV w fp8 dla rzadkiej uwagi na sm86. Efekt: pula 806 792 tokeny, czyli trzy pełne sesje po 262 tys. albo pięć po ~160 tys. naraz.
Łatki to 58 commitów od autora z kontem założonym w lipcu 2026 — dlatego zanim cokolwiek trafiło na serwer, pięciu agentów równolegle przeczytało 15,5 tys. linii kodu, który działa w produkcji: jądra CUDA, transport tablicy n-gramów, telemetrię, silnik oraz skrypty i 196 zależności. Wynik: żadnej komunikacji z siecią, żadnego czytania kluczy, żadnego zapisu treści promptów. Sprawdziliśmy też, że paczka to dokładnie oficjalny v0.30.0 plus te commity (zgodny hash drzewa). Dwa moduły CUDA kompilujemy sami (18 min), resztę bierzemy z oficjalnego pakietu vLLM z PyPI — gotowych binarek autora nie używamy. Jego skrypty instalacyjne tworzyły uprzywilejowany kontener i wystawiały API bez klucza, więc też poszły w odstawkę.
Bufor pierścieniowy rzadkiej uwagi nie jest zerowany, gdy blok pamięci trafia do nowego zapytania — w oficjalnym v0.30.0 nowe zapytanie mogło więc teoretycznie natrafić na bajty poprzedniego. Łatka dopisuje do każdego wiersza znacznik pozycji i odrzuca wiersze, które do zapytania nie należą. Działa to tylko przy tzw. fused pre-indexerze — z konfiguracji modelu sprawdziliśmy, że używają go wszystkie warstwy. Kod od społeczności bywa bezpieczniejszy od oficjalnego — ale trzeba to sprawdzić, a nie założyć.
51 mld parametrów tablicy n-gramów nie musi siedzieć w VRAM — na token czyta się z niej kilka wierszy. Oficjalny vLLM trzyma osobną kopię dla każdej karty (4 × 48 GB), łatka obsługuje ją w osobnym procesie, w jednej kopii (~50 GB z 256 GB RAM). Ten proces rozmawia z resztą przez gniazdo uniksowe z pickle, więc usługa dostaje prywatny katalog 0700, umask 077 i sprzątanie /dev/shm przy każdym starcie — po awarii zostają tam tokeny ostatniego zapytania.
Dwie pary kart w potoku z równoległością ekspertów, każda karta ~22,5 GB z 24. Pierwszy start trwa ok. 12 minut — ładowanie 115 GiB wag i kompilacja grafów CUDA. MTP z trzema tokenami szkicu, rozumowanie domyślnie w trybie „low”, a klucz API, wyłączona telemetria vLLM i automatyczny restart są w usłudze systemd, nie w skrypcie autora.
Przy trzech osobach pierwsza powtórka dała 118 tok/s, druga 189; przy pięciu 155 i 239. FlashInfer kompiluje jądra przy pierwszym napotkanym rozmiarze wsadu. W tabelach podajemy drugą powtórkę. Mierz zawsze co najmniej dwa razy i nie bierz pierwszego przebiegu po starcie za wynik.
Przy dokumencie 47 tys. tokenów Flash-Next generuje 134 tok/s, przy 189 tys. — 140 tok/s. Qwen3.8-27B na tych samych kartach spada z 59 do 26, a potem do 9 tok/s. Czekanie na pierwszy token nadal rośnie z długością dokumentu (15 s → 64 s), bo tekst trzeba przeczytać — ale czyta go 2–2,5 raza szybciej niż 27B.
Ten sam stresstest co dla 27B, z sufitem rozmów podniesionym do 95 tys.: 62 minuty, 339 rund, rozmowy do 107 tys. tokenów, zero błędów, zero restartów, zero Xid/AER, karty na 220 W, najgorętsza 79 °C. Przy tej samej długości rozmowy Flash-Next generuje 2–2,7 raza szybciej (45–60 tys.: 2,9 → 7,4 tok/s), ale liczby bezwzględne nadal są niskie: długie prompty ośmiu osób zjadają moc, którą karty dałyby na generowanie, a PCIe Gen3 bez NVLinka dokłada swoje.
Jak bardzo 4 bity szkodzą jakości odpowiedzi względem 27B w pełnej precyzji — porównanie dokładności jest w toku. Skale KV fp8 pochodzą od autora łatek; jeśli będziemy kalibrować sami, to tylko na korpusie bez danych klientów.
Flash-Next nie jest jednym wynalazkiem, tylko złożeniem kilku idei z ostatnich trzech lat. Każda z nich rozwiązuje inny problem: jak mieć dużo wiedzy bez dużego liczenia, jak czytać długi tekst bez spowalniania i co trzymać w drogiej pamięci karty, a co w taniej pamięci komputera.
| Kiedy | Co | Co z tego wynikło |
|---|---|---|
| 12.2023 | Mixtral 8×7B (Mistral) | Mieszanka ekspertów w otwartym modelu: dużo parametrów, mało liczenia na token. |
| 12.2023 | Mamba | Warstwy z pamięcią o stałym rozmiarze zamiast uwagi, której koszt rośnie z długością tekstu. |
| 2024 | DeepSeek-V2 i V3 | Wielu drobnych ekspertów plus ekspert wspólny; w V3 przewidywanie kilku tokenów naraz (MTP), które dziś przyspiesza generowanie. |
| 2024–2025 | DeltaNet → Gated DeltaNet | Liniowa uwaga, która potrafi nadpisywać i wygaszać pamięć — dużo lepsza w wyszukiwaniu w tekście niż wcześniejsze warianty. |
| 2025 | Gemma 3n (Google) | Osadzenia per warstwa trzymane poza akceleratorem — wzorzec dla tablicy n-gramów w RAM. |
| 09.2025 | Qwen3-Next 80B-A3B | Pierwszy hybrydowy Qwen: trzy warstwy Gated DeltaNet na jedną warstwę pełnej uwagi, 512 ekspertów, MTP. |
| 09.2025 | DeepSeek-V3.2-Exp | Rzadka uwaga z lekkim indekserem, który wybiera, które fragmenty tekstu w ogóle czytać. |
| 2026 | Qwen3.8-Flash-Next | Wszystko naraz: hybryda 3:1, rzadka uwaga QSA w co czwartej warstwie, 51 mld parametrów w tablicy n-gramów, MTP. 177 mld parametrów, ok. 6 mld aktywnych. |
| 31.08.2026 | vLLM | Obsługa modelu w głównej gałęzi; wydanie v0.30.0 — 22.09.2026. Na Ampere z KV w fp8 — tylko z łatkami społeczności. |
| Kiedy | Co | Jeden użytkownik | Ośmiu — łącznie |
|---|---|---|---|
| 08.2026 | Twarde resety pod obciążeniem: jedna karta w gnieździe na chipsecie, inna z łączem x2 zamiast x8 przez riser. Naprawione przełożeniem kart. | — | — |
| 25.09.2026 | Qwen3.8-27B BF16, bez MTP → z MTP (2 tokeny) | 40,3 → 55,5 tok/s | 140,0 → 148,9 tok/s |
| 04.10.2026 | Qwen3.8-Flash-Next 177B INT4, KV fp8, MTP (3 tokeny) | 120,6 tok/s | 290,7 tok/s |
Do 25 września 2026 roku B70 pracował u nas na llama.cpp SYCL — na Intelu to przez długi czas była jedyna droga, która działała stabilnie. Na vLLM dla kart Arc Pro czekaliśmy, aż dojrzeje: Intel wydaje go w projekcie llm-scaler, obsługę MTP dla Qwen 27B dodał w wersji z 4 sierpnia 2026, a wersję intel/llm-scaler-vllm:0.26.0-b2 z poprawkami cache prefiksu i współbieżności — 9 września. Do tego społeczność opublikowała kwantyzację GPTQ int4 Qwen3.8-27B z zachowaną wizją i warstwą MTP. 25.09.2026 zmierzyliśmy vLLM na tej samej karcie, a 26.09.2026 przełączyliśmy na niego model. Klienci nie zauważyli zmiany — wszyscy łączą się przez bramę API pod stałą nazwą modelu, więc backend podmienia się w jednym miejscu.
Te same testy, ta sama karta, ten sam limit 250 W, prompt 2 849 tokenów, 600 tokenów odpowiedzi:
| llama.cpp SYCL (do 26.09) | vLLM XPU bez grafów | vLLM XPU grafy + cache — teraz | |
|---|---|---|---|
| 1 użytkownik — generowanie | 20,1 tok/s | 11,8 tok/s | 32,8 tok/s |
| 1 użytkownik — łącznie z promptem | 17,8 tok/s | 11,5 tok/s | 30,3 tok/s |
| Czas do 1. tokenu / czytanie promptu | 3,9 s / 735 tok/s | 1,4 s / 1 975 tok/s | 1,5 s / 1 917 tok/s |
| 3 użytkowników — łącznie | 26,7 tok/s | 36,8 tok/s | 75,5 tok/s |
| 8 użytkowników — łącznie | — (3 sloty) | 84,3 tok/s | 144,3 tok/s |
| Generowanie przy 15k / 60k / 120k kontekstu | 17,8 / 12,0 / 8,4 | 12,1 / 12,0 / 12,0 | 31,5 / 28,0 / 24,5 |
| Czytanie promptu przy 15k / 60k / 120k | 937 / 814 / 693 | 1 789 / 1 273 / 918 | 1 739 / 1 273 / 919 |
| Powrót do rozmowy (cache prefiksu) | 1,5–2,6 s | 8–131 s (brak cache) | 0,5–1,2 s |
| Stress: 2 osoby × 30 min, do 122k | 40 rund, 0 błędów | — | 49 rund, 0 błędów |
| Stress: 8 osób × 60 min, do 61k | — | — | 160 rund, 0 błędów |
| Kwantyzacja / cache KV | GGUF UD-Q4_K_M / q8_0 | GPTQ int4 / fp8 | GPTQ int4 / fp8 |
| Kontekst | 3 × 131 072 | 131 072 | 131 072 na rozmowę, pula wspólna |
--enforce-eager — kosztuje ×2,8Wszystkie przepisy w dokumentacji llm-scaler uruchamiają vLLM z --enforce-eager. Z nim jeden użytkownik dostaje 11,8 tok/s — wolniej niż na llama.cpp. Grafy XPU włącza zmienna VLLM_XPU_ENABLE_XPU_GRAPH=1 (bez --enforce-eager): 32,8 tok/s. Kompilacja grafu przy pierwszym starcie trwa ~150 s — trzymaj cache kompilacji na trwałym dysku, to ten sam błąd co --enforce-eager na 3090 (punkt wyżej), tylko w innej zmiennej.
--enable-prefix-cachingQwen3.8 to model hybrydowy (warstwy Gated DeltaNet + attention). W tej wersji vLLM bez jawnego --enable-prefix-caching cache prefiksu jest wyłączony — powrót do rozmowy na 120 tys. tokenów oznaczał 131 s czytania od nowa. Z flagą vLLM przechodzi w tryb mamba_cache_mode 'align' (oznaczony jako eksperymentalny) i powrót trwa 0,5–1,2 s.
--gpu-memory-utilizationBudżet pamięci vLLM obejmuje wagi i cache KV, a grafy XPU (2,84 GiB) dochodzą ponad niego. Zajętość rośnie jeszcze w trakcie pracy, gdy pula KV się zapełnia i przy pierwszym obrazie. Przy 0.90 po długim kontekście karta miała zajęte 31,78 z 31,89 GiB. Stąd niższy budżet niż na NVIDII (tam 0.94) — i powód jest poważniejszy niż wydajność:
Gdy na B70 kończy się VRAM, sterownik xe (TTM) wypycha bufory karty do RAM-u hosta, zamiast zwrócić błąd alokacji. To pamięć jądra — limit RAM-u kontenera jej nie obejmuje. U nas 16 GB hosta zniknęło w minutę: [TTM] Buffer eviction failed, globalny OOM, host bez SSH przez 23 minuty. Zmiany, które dokładają VRAM, sprawdzaj od bezpiecznej wartości w górę, z zapasem ≥ 1 GiB po pierwszym obrazie (bufor wizji alokuje się dopiero wtedy, +0,5 GiB), i ze strażnikiem, który zabije serwer przy pierwszym wpisie TTM w dmesg.
Spekulacja MTP (qwen3_5_mtp, 2 tokeny) dała jednemu użytkownikowi ~51 tok/s generowania. Ale dokłada ~0,8 GiB wag i powiększa grafy z 2,8 do 4,75 GiB — przy budżecie 0.90 karta była pełna już po starcie (31,84 GiB). Przy trzech osobach jeden przebieg zawisł na 31 s, przy ośmiu VRAM się przepełnił. Na serwerze dla wielu osób — bez MTP. Wariant dla ≤ 5 osób z niższym budżetem sprawdzimy osobno.
Przy ośmiu osobach wklejających po ~50 stron co rundę pula KV (191 tys. tokenów) się przepełnia i vLLM odkłada zapytania, licząc je potem od nowa — mediana czasu do pierwszego tokenu 92 s, pojedyncze do 4,5 minuty. Zero błędów, ale przy takim ruchu to czuć. Z820 z pulą 431 tys. w tym samym teście nie wywłaszczał ani razu. Kontekst jednej rozmowy (131 072) od budżetu nie zależy — budżet decyduje, ile długich rozmów zmieści się naraz.
Build llama.cpp SYCL jest wrażliwy na wersję oneAPI i runtime'u. Aktualizacja „przy okazji” potrafi zepsuć działający stack. Blokujemy wersje pakietów i aktualizujemy świadomie.
--parallel dzieli kontekstW llama.cpp -c to kontekst całkowity, dzielony na sloty. -c 393216
--parallel 3 daje 3 × 131 072, a nie 3 × 393 216. KV w q8_0:
262k kontekstu ≈ 11,5 GiB ponad wagi.
-b 4096 -ub 1024 zamiast domyślnych 2048/512: czytanie promptu 579 → 716 tok/s (+24 %), czas do pierwszego tokenu −19 %. Generowanie bez zmian — i tak ma być.
--spec-type ngram-cacheSerwer wstaje w 150 s i pada przy pierwszym zapytaniu. Sześć prób, zawsze „connection refused”, dziennik pusty. Pozostałe warianty n-gramowe czekają na sprawdzenie.
Pod obciążeniem 237 W przy limicie 250 W. Podniesienie do 275 W: +0,2 % dla jednego użytkownika, +4 % dla trzech. Za 20 W i głośniejszy wentylator — nie warto.
Nowsze llama.cpp trzyma zapisane rozmowy w RAM-ie hosta — --cache-ram, domyślnie 8192 MiB. Stan jednej rozmowy na ~120 tys. tokenów z KV q8_0 to u nas 5,9 GB. Kontener z 8 GB RAM-u padał na tym po 13 minutach stress testu (OOM killer), a przy mniejszym kontekście zawisał na 20 minut w swapie. Poprawka: --cache-ram poniżej limitu pamięci kontenera (u nas 2048). Po zmianie: 30 minut, dwóch użytkowników, rozmowy do 122 tys. tokenów — bez restartu, RAM maks. 5,3 GB.
GGUF Qwen3.8-27B ma warstwę MTP (blk.64.nextn.*), a llama.cpp ma --spec-type draft-mtp. Jeden użytkownik: 20,1 → 26,9 tok/s. Trzech naraz: 26,7 → 14,3 tok/s łącznie — szkic liczony jest osobno dla każdego slotu. Do tego +2,6 GiB VRAM. Na serwerze dla kilku osób — nie.
Czysty pomiar, jeden użytkownik, nikt inny na serwerze: przy 15 / 60 / 120 tys. tokenów kontekstu generowanie 17,8 / 12,0 / 8,4 tok/s, czytanie promptu 937 / 814 / 693 tok/s. Powrót do tej samej rozmowy z cache — 1,5–2,6 s do pierwszego tokenu. Szybka ścieżka flash attention dla Battlemage (llama.cpp PR #25222) działa tylko przy KV f16 i jednej sekwencji: KV f16 dało u nas +16 % generowania przy 60 tys., ale kosztuje połowę kontekstu.
Liczyliśmy tokeny po polu reasoning_content, a vLLM 0.20.1 wysyła reasoning. Zero rozpoznanych tokenów, TTFT równy całej odpowiedzi, „6,66 tok/s” zamiast prawdziwych 8,4. Zawsze wypisz, ile kawałków strumienia niesie treść. Zero = mierzysz szum.
max() z powtórzeń odwraca rankingRozrzut sięgał 18 % (45,5 i 54,5 tok/s w tej samej konfiguracji). Najlepszy przebieg zamiast średniej potrafił postawić słabszą maszynę przed mocniejszą. Zawsze średnia, zawsze z rozrzutem.
nvidia-smi mierzy nie tę kartęW maszynie z B70 i RTX 3060 skrypt raportował 15 W i 210 MHz — bezczynną 3060. Telemetria Intela jest w /sys/class/hwmon/*/ z name = xe: energy1_input (µJ), power1_cap, temp*_input.
memory.usedSkrypt porównywał licznik restartów usługi przed i po teście. Serwer 20 minut wisiał w swapie, ale się nie zrestartował — więc „przetrwał cały test”, a dwie ostatnie rundy wróciły puste (0 tokenów, 1 187 s do pierwszego tokenu) zamiast z błędem. Sprawdzaj liczbę tokenów w odpowiedzi, czas do pierwszego tokenu i zdarzenia OOM w cgroupie, nie tylko restarty.
Gdy jeden użytkownik wkleja 15 tys. tokenów, jego prompt idzie w tych samych krokach co generowanie drugiego — ten dostaje token co kilka sekund. Stąd „4 tok/s” przy dwóch osobach na B70 i 2,4–3,2 tok/s przy ośmiu na Z820. To nie jest spowolnienie od długiego kontekstu — do tego potrzebny jest osobny pomiar jednego użytkownika.
Średnia moc z różnicy licznika na początku i na końcu 30-minutowego testu wyszła 93 W, a próbkowanie co sekundę w trakcie dawało 242 W. Czytaj licznik często i sumuj przyrosty.
| Format / repo | Werdykt dla RTX 3090 (Ampere) i Arc B70 |
|---|---|
NVFP4 (unsloth/…-NVFP4 i in.) | tylko Blackwell (RTX 50xx) — odpada |
FP8 (Qwen/…-FP8) | Ada lub Hopper — na Ampere odpada |
| NVFP4-MTP-GGUF | ma głowicę MTP, ale NVFP4 — dla Intela bezużyteczny |
| MLX | format Apple Silicon |
| int4 AutoRound / AWQ-INT4 | działa na 3090 (ścieżka Marlin) — tego używamy |
GGUF UD-Q4_K_M / ggml-org | działa na B70 przez llama.cpp SYCL (do 26.09 nasz wybór) |
GPTQ int4 + MTP (SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16) | działa na B70 przez vLLM XPU (intel/llm-scaler-vllm), wizja w komplecie — tego używamy od 26.09 |
Na 3090 żaden nowy format nie pomoże. Na B70 pomógł nie format, tylko silnik — ten sam model w GPTQ int4 pod vLLM XPU jest 1,6–2,9× szybszy niż GGUF pod llama.cpp. Zanim pobierzesz 20 GB, sprawdź, czy Twoja architektura w ogóle go obsługuje.
Same model (Qwen3.8-27B), same method: 2,849-token prompt, 600-token answer, reasoning on, means of repeated runs. Lighter bar — before the change, solid bar — after. Full tables are in the benchmark; what exactly we changed is in the points below.
--enforce-eager (CUDA graphs), MTP speculation and a larger queue took decode from 8.4 to 28.9 tok/s — ×3.4 without touching the hardware. Another +45 % came from a 220 W cap instead of 180 W on a machine whose PSU can take it (Z440: 32.3 → 46.8 tok/s).llm-scaler): XPU graphs landed in July 2026 as experimental and are still off in the examples, MTP for Qwen 27B arrived on 4 August, prefix-cache and concurrency fixes on 9 September, and MTP with many users still breaks. Our B70 ran from 2 September; for the first three weeks llama.cpp was the only stable route. Buy Intel and you get cheaper memory right away — the performance arrives later, with the software.Measured for the GPU and CPU package only, from driver energy counters and RAPL. It excludes the motherboard, drives, case fans, VRM losses and power-supply efficiency — actual draw at the wall is higher.
Same machine, same card, same 180 W power limit and the same 1500 MHz clock lock. Only the vLLM engine configuration changed. Measured on the Z420 on 2026-09-24, means over two repeats of every run.
Same machine, same vLLM configuration, only the GPU power limit changed. The 1500 MHz clock lock stays on throughout.
--max-num-seqs 4 and
--gpu-memory-utilization 0.94. The KV cache pool is then 111,287 tokens,
i.e. 1.7 full-length requests at once, or correspondingly more shorter ones.--max-num-seqs 8), because those settings let the engine serve more
requests at once. Single-user speed is identical in both variants — 28.9 tok/s at 32k and 28.9 tok/s
at 64k — so the single-user rows hold for both. What drives memory use is the queue size rather than
the window itself: CUDA graphs take more the more batch sizes they have to capture.llm-scaler-vllm:0.26.0-b2 image, GPTQ int4, fp8 KV cache, XPU graphs, prefix cache), measured with the same prompt (2,849 tok.) and method as the Z820 and Z440 at a 250 W cap. The earlier llama.cpp figures are in the B70 section below, as history. Idle power still comes from 2026-09-10.--enforce-eager removed (so CUDA graphs on),
MTP speculation kept, the --max-num-seqs queue raised from 3 to 8,
--enable-chunked-prefill enabled, and the context window shortened to 32k tokens so
that there is memory left for the graphs. No hardware change at all.--enforce-eager the host has to launch every compute kernel by hand.
With graphs enabled the card runs at its cap and becomes the bottleneck itself, rather than the
machine. Hence a threefold jump without replacing anything.Until 4 October 2026 the fourth benchmark column was a single RTX 3090 in an HP Z440 (Xeon E5-2683 v3, 62 GB DDR4). Four 3090s with Flash-Next replaced it; the results stay here because it is still the cheapest way to a local model. Qwen3.8-27B int4 AutoRound, fp8 KV, 65,536 window, max-num-seqs 8, vLLM 0.20.1, measured 2026-09-24.
| 3090 · Z440 · 180 W | |
|---|---|
| Single user — wall time / tok/s | 22.7 s / 26.4 |
| Decode only | 32.3 tok/s |
| Time to first token / prompt processing | 4.2 s / 675 tok/s |
| Three users — total / per user | 48.5 / 16.2 tok/s |
| Eight users — total | 45.8 tok/s |
| GPU / CPU / combined draw | 179 W / 43 W / ~222 W |
| Electricity under load — measured / at the wall | €39.96 / €59.76 per month |
| Hardware per 1 tok/s (1 / 3 users) | €40 / €22 |
| Throughput per watt (1 / 3 users) | 0.147 / 0.271 tok/s/W |
| Power limit | Decode single user | Single user tok/s | 3 at once aggregate | 8 at once aggregate | Time to first token | Prompt processing | Temp. |
|---|---|---|---|---|---|---|---|
| 180 W | 32.3 tok/s | 26.4 | 48.5 tok/s | 45.8 tok/s | 4.2 s | 675 tok/s | 74 °C |
| 200 W | 38.5 tok/s | 31.3 | 60.1 tok/s | 57.7 tok/s | 3.6 s | 790 tok/s | 79 °C |
| 220 W | 46.8 tok/s | 37.7 | 71.9 tok/s | 69.1 tok/s | 3.1 s | 919 tok/s | 82 °C |
Both engines default to processing a single sequence at a time. This is not a hardware limit — it is one configuration value, and raising it costs no additional VRAM, because the KV pool is allocated once at startup.
| Engine | Parameter | Change | Throughput, three users |
|---|---|---|---|
| vLLM | --max-num-seqs | 3 → 24 | 33.3 → 63.7 tok/s at eight users (×1.91) |
| llama.cpp | --parallel | 1 → 3 | 18.95 → 27.05 tok/s (×1.43) |
The queue only helps when speculation is off, though. With MTP speculation enabled, raising the queue from 3 to 8 changed almost nothing (45.6 → 46.1 tok/s at eight users) — speculation consumes the same spare compute that batching lives on.
Why the Arc scales less well (×1.44 against ×1.91): at a single stream the Arc already saturates its compute — per-user throughput falls from 19 to 10 tok/s with three. The RTX 3090 at a single stream is memory-bandwidth bound and has compute to spare. Counter-intuitively, the faster a card is on a single stream, the less it gains from batching.
With reasoning enabled the model spends roughly 350 tokens on the thinking block alone. At a 600-token limit the answer comes back randomly empty — 996 characters of content on one run, zero on the next, with identical settings. That is variance at the edge of the budget, not a fault.
At 2,500 tokens: 1,407 characters of reasoning and 6,986 characters of actual answer. The practical floor is 1,500–2,000 tokens.
The 3090 runs with a 180 W power limit. Adding a 1,500 MHz clock lock — even though the card is capable of 1,980 MHz — raised single-user throughput from 6.8 to 8.0 tok/s.
Under a power limit alone the card boosts to its maximum, immediately hits the power ceiling and is throttled hard, producing a sawtooth. Pinned to a lower fixed clock it stays inside the power budget continuously. A stable lower clock beats an oscillating higher one — and it also flattens the microsecond current spikes that a power limit alone does nothing about.
On the Arc, a 262,144-token context with a q8_0 KV cache occupies about 11.5 GiB on top of the model weights themselves (14.3 GiB). Raising the total context to 393,216 and splitting it across three slots gives each user 131,072 tokens:
before: 25.76 GiB / 32 1 slot × 262,144 after: 30.65 GiB / 32 3 slots × 131,072 no allocation errors
On vLLM the relationship runs the other way: enabling CUDA graphs cut
the available KV cache from 4.61 to 2.24 GiB, while a 64k context needs 2.3 GiB — the engine
refused to start. With speculative decoding enabled the graphs fall back to partial mode
(PIECEWISE), because the FlashInfer backend does not support the full one.
Partial mode still pays off handsomely: 13.3 → 31.7 tok/s per user on the
same machine at the same power limit. MTP speculation accounts for ×1.7 of that — without it
the same settings give 19.0 tok/s. The trap is in the memory accounting: graphs allocate
outside the --gpu-memory-utilization budget, so it has to be
lowered or the context shortened. With a queue of 24 and speculation on, no setting was found
where the engine starts at all on 24 GB.
The xe driver publishes what GeForce cards do not: separate energy counters
for the board and the package, temperatures for the core, memory, memory controller and the
PCIe interface, plus sixteen individual VRAM channels, fan speed, and both
the power cap and the critical threshold.
The RTX 3090 reports no voltages at all — a deliberate restriction on consumer cards — and no temperatures beyond the core. The only substitute is a power-braking counter that increments when the card is starved of power.
--enforce-eager can cost you ×3Without CUDA graphs the host launches every kernel by hand. On an old Xeon the card waited on the CPU and drew 131–138 W against a 180 W limit. Flag removed: 8.4 → 28.9 tok/s on the same machine.
Tell-tale sign: the card never reaches its own power limit under load. The host is the bottleneck, not the GPU.
--gpu-memory-utilizationEnabling graphs cut the available KV cache from 4.61 to 2.24 GiB and the engine refused to
start at 64k context. The larger --max-num-seqs, the more batch sizes to capture
and the more memory they take.
# works on 1× 3090, 64k, KV pool 111,287 tok. --max-num-seqs 4 --gpu-memory-utilization 0.94 --max-model-len 65536
With a queue of 24 plus speculation we found no setting that starts on 24 GB at all.
MTP gives ×1.7 for a single user. But with MTP on, raising the queue from 3 to 8 did nothing (45.6 → 46.1 tok/s at 8 users). Pick for your traffic: single stream → MTP, many users at once → queue without speculation.
With MTP the graphs fall back to PIECEWISE (FlashInfer lacks full mode).
Still worth it.
Prefix caching makes a second run with the same prompt "magically" faster. Put a unique marker in every request. Measure with streaming to separate time to first token from decode.
max_tokens = randomly empty answersThe thinking block is ~350 tokens. At a 600 limit the answer is 996 characters one run and zero the next, same config. Practical minimum: 1,500–2,000 tokens.
--gpu-memory-utilization 0.96 starts, then diesThe engine comes up fine, then under load: torch.OutOfMemoryError: Tried to allocate 56.00 MiB. The sneakiest one, because startup looks like success. Stay at 0.95 or below. Raising the budget also shrinks the room for graphs — the real lever is --max-num-seqs.
--enable-chunked-prefill is on by default in vLLM 0.20.1 — the log shows enable_chunked_prefill=True without the flag--kv-cache-dtype takes only auto and fp8 variants. llama.cpp has 4-bit KVThe startup log, non-default args line, contains 'api_key': [...]. Anyone reading the journal has the key. Rotate it before sharing logs or the machine.
Same 3090 on DDR3 and DDR4: 35 % apart with --enforce-eager, 12 % with graphs on. Don't replace the board or RAM before checking the engine config.
With only a 180 W limit the card boosts, hits the ceiling and gets throttled. Locking it at 1500 MHz: 6.8 → 8.0 tok/s despite the lower clock. It also removes some of the power spikes the limit cannot catch.
nvidia-smi -pl 180 nvidia-smi -lgc 1500,1500
180 → 220 W gave +45 % decode, but only after graphs were on. In eager mode changing the limit did nothing, because the card never reached 180 W anyway.
At 250 W in an old HP chassis the PSU's OCP tripped under combined CPU and GPU load. The machine just goes dark, no logs. A power limit caps the average over milliseconds, not the microsecond spikes during batching — those trip the OCP. Another such machine died at 200 W. The power limit belongs to the chassis PSU, not the card — raise it 20 W at a time.
--speculative-config '{"method": "mtp", "num_speculative_tokens": 2}' on Qwen3.8-27B in BF16, tensor-parallel 4: single-user decode 40.3 → 55.5 tok/s. At three and eight users the gain shrinks to 4–6 % (79.6 → 82.4 and 140.0 → 148.9 tok/s aggregate), because the GPUs already run full batches. Draft acceptance over a one-hour test: 73 % — 82 % at position one, 64 % at position two.
--disable-custom-all-reduce does nothing on four PCIe GPUsvLLM prints at start-up: Custom allreduce is disabled because it's not supported on more than two PCIe-only GPUs. Without NVLink and with more than two cards custom all-reduce is always off. We compared “with the flag” and “without” and got 136 against 162 tok/s at eight users — that was one weak run out of four, not the flag. Before crediting a change with a difference, check the engine log to see whether the change took effect at all.
Eight users for 60 minutes, each pasting a long document every round, conversations up to 50k tokens: 215 rounds, zero errors, zero restarts, zero Xid/AER, GPUs at their 220 W cap, hottest at 81 °C. But the prefix cache hit only 40 % of the time — eight conversations of ~50k tokens fill almost the whole KV pool (431k in BF16), so one conversation's blocks get evicted by other users' requests before the same user comes back. vLLM then re-reads the whole conversation (~2,200 tok/s on four GPUs) and time to first token climbs to ~33 s. The lever is an fp8 KV cache — twice the pool at some cost in quality.
Same Z820, new model. Flash-Next has 125 billion main parameters, 51 billion in an n-gram table and an MTP head. Each expert layer has 512 experts, of which 10 plus one shared expert run per token — about 6 billion active parameters in total. Three layers out of four are Gated DeltaNet with a fixed-size memory; every fourth is Qwen Sparse Attention, which reads only selected blocks of text. 262,144 tokens natively.
Ampere (sm86) has no hardware FP8. Official vLLM supports this model since 0.30.0 (2026-09-22), but on a 3090 only with a 16-bit conversation memory (KV) — which gives a pool of about 400k tokens. The way in is INT4 W4A16 quantisation on Marlin kernels plus community patches that add an fp8 KV reader for sparse attention on sm86. Result: an 806,792-token pool — three full 262k sessions or five of ~160k at once.
The patches are 58 commits from an author whose account dates from July 2026 — so before anything reached the server, five agents read the 15.5 thousand lines of code that run in production in parallel: CUDA kernels, the n-gram table transport, telemetry, the engine, plus the scripts and 196 dependencies. Result: no network traffic, no key reading, no prompt logging. We also checked that the bundle is exactly official v0.30.0 plus those commits (matching tree hash). We compile the two CUDA modules ourselves (18 min) and take the rest from the official vLLM package on PyPI — no binaries from the author. His install scripts created a privileged container and exposed the API without a key, so they were dropped as well.
The sparse-attention ring buffer is not zeroed when a memory block goes to a new request — so in official v0.30.0 a new request could in theory meet the previous one's bytes. The patch tags every row with its position and rejects rows that do not belong to the request. This only works with the so-called fused pre-indexer — we checked from the model configuration that every layer uses it. Community code can be safer than the official one — but you have to verify that, not assume it.
The 51-billion-parameter n-gram table does not need to live in VRAM — only a few rows are read per token. Official vLLM keeps a separate copy per GPU (4 × 48 GB); the patch serves it from a separate process, in one copy (~50 GB of 256 GB RAM). That process talks to the rest over a Unix socket with pickle, so the service gets a private 0700 directory, umask 077 and a /dev/shm clean-up on every start — after a crash the last request's tokens are left there.
Two pairs of GPUs in a pipeline with expert parallelism, each card at ~22.5 GB of 24. The first start takes about 12 minutes — loading 115 GiB of weights and compiling CUDA graphs. MTP with three draft tokens, reasoning in “low” mode by default; the API key, disabled vLLM telemetry and automatic restart live in the systemd service, not in the author's script.
At three users the first run gave 118 tok/s and the second 189; at five, 155 and 239. FlashInfer compiles kernels the first time it meets a batch size. The tables show the second run. Always measure at least twice and never take the first pass after a start as the result.
On a 47k-token document Flash-Next generates 134 tok/s; at 189k, 140 tok/s. Qwen3.8-27B on the same cards drops from 59 to 26 and then to 9 tok/s. The wait for the first token still grows with document length (15 s → 64 s), because the text has to be read — but it is read 2–2.5 times faster than with 27B.
The same stress test as for 27B, with the conversation ceiling raised to 95k: 62 minutes, 339 rounds, conversations up to 107k tokens, zero errors, zero restarts, zero Xid/AER, cards at 220 W, the hottest at 79 °C. At the same conversation length Flash-Next generates 2–2.7 times faster (45–60k: 2.9 → 7.4 tok/s), but the absolute numbers are still low: eight users' long prompts eat the compute the cards would spend on generation, and PCIe Gen3 without NVLink adds its share.
How much 4-bit hurts answer quality against 27B at full precision — the accuracy comparison is in progress. The fp8 KV scales come from the patch author; if we calibrate ourselves, it will only be on a corpus with no client data.
Flash-Next is not a single invention but an assembly of several ideas from the last three years. Each solves a different problem: how to hold a lot of knowledge without a lot of compute, how to read long text without slowing down, and what to keep in expensive GPU memory versus cheap system memory.
| When | What | What came of it |
|---|---|---|
| 12.2023 | Mixtral 8×7B (Mistral) | Mixture of experts in an open model: many parameters, little compute per token. |
| 12.2023 | Mamba | Layers with a fixed-size memory instead of attention whose cost grows with text length. |
| 2024 | DeepSeek-V2 and V3 | Many fine-grained experts plus a shared expert; in V3, multi-token prediction (MTP), which now speeds up generation. |
| 2024–2025 | DeltaNet → Gated DeltaNet | Linear attention that can overwrite and decay its memory — much better at retrieval from text than earlier variants. |
| 2025 | Gemma 3n (Google) | Per-layer embeddings kept off the accelerator — the pattern behind the n-gram table in RAM. |
| 09.2025 | Qwen3-Next 80B-A3B | The first hybrid Qwen: three Gated DeltaNet layers per full-attention layer, 512 experts, MTP. |
| 09.2025 | DeepSeek-V3.2-Exp | Sparse attention with a lightweight indexer that picks which parts of the text to read at all. |
| 2026 | Qwen3.8-Flash-Next | All of it at once: a 3:1 hybrid, QSA sparse attention in every fourth layer, 51 billion parameters in an n-gram table, MTP. 177 billion parameters, about 6 billion active. |
| 2026-08-31 | vLLM | Model support in the main branch; release v0.30.0 on 2026-09-22. On Ampere with fp8 KV — only with community patches. |
| When | What | Single user | Eight — total |
|---|---|---|---|
| 08.2026 | Hard resets under load: one card in a chipset slot, another linked at x2 instead of x8 through a riser. Fixed by moving the cards. | — | — |
| 2026-09-25 | Qwen3.8-27B BF16, without MTP → with MTP (2 tokens) | 40.3 → 55.5 tok/s | 140.0 → 148.9 tok/s |
| 2026-10-04 | Qwen3.8-Flash-Next 177B INT4, fp8 KV, MTP (3 tokens) | 120.6 tok/s | 290.7 tok/s |
Until 25 September 2026 our B70 ran on llama.cpp SYCL — for a long time the only route on Intel that ran reliably. We waited for vLLM on Arc Pro to mature: Intel ships it in the llm-scaler project, added MTP support for Qwen 27B in the release of 4 August 2026, and published intel/llm-scaler-vllm:0.26.0-b2 with prefix-cache and concurrency fixes on 9 September. The community also published a GPTQ int4 quant of Qwen3.8-27B that keeps vision and the MTP layer. On 2026-09-25 we measured vLLM on the same card, and on 2026-09-26 we switched the model over to it. Clients did not notice — they all connect through the API gateway under a fixed model name, so the backend is swapped in one place.
Same tests, same card, same 250 W cap, 2,849-token prompt, 600-token answer:
| llama.cpp SYCL (until 09-26) | vLLM XPU no graphs | vLLM XPU graphs + cache — now | |
|---|---|---|---|
| Single user — decode | 20.1 tok/s | 11.8 tok/s | 32.8 tok/s |
| Single user — incl. prompt | 17.8 tok/s | 11.5 tok/s | 30.3 tok/s |
| Time to first token / prompt processing | 3.9 s / 735 tok/s | 1.4 s / 1,975 tok/s | 1.5 s / 1,917 tok/s |
| Three users — aggregate | 26.7 tok/s | 36.8 tok/s | 75.5 tok/s |
| Eight users — aggregate | — (3 slots) | 84.3 tok/s | 144.3 tok/s |
| Decode at 15k / 60k / 120k of context | 17.8 / 12.0 / 8.4 | 12.1 / 12.0 / 12.0 | 31.5 / 28.0 / 24.5 |
| Prompt processing at 15k / 60k / 120k | 937 / 814 / 693 | 1,789 / 1,273 / 918 | 1,739 / 1,273 / 919 |
| Returning to a conversation (prefix cache) | 1.5–2.6 s | 8–131 s (no cache) | 0.5–1.2 s |
| Stress: 2 users × 30 min, up to 122k | 40 rounds, 0 errors | — | 49 rounds, 0 errors |
| Stress: 8 users × 60 min, up to 61k | — | — | 160 rounds, 0 errors |
| Quantisation / KV cache | GGUF UD-Q4_K_M / q8_0 | GPTQ int4 / fp8 | GPTQ int4 / fp8 |
| Context | 3 × 131,072 | 131,072 | 131,072 per conversation, shared pool |
--enforce-eager — it costs ×2.8Every recipe in the llm-scaler docs starts vLLM with --enforce-eager. With it a single user gets 11.8 tok/s — slower than llama.cpp. XPU graphs are enabled with VLLM_XPU_ENABLE_XPU_GRAPH=1 (no --enforce-eager): 32.8 tok/s. Graph compilation on first start takes ~150 s — keep the compile cache on persistent storage. It is the same mistake as --enforce-eager on the 3090 (above), just behind a different switch.
--enable-prefix-cachingQwen3.8 is a hybrid model (Gated DeltaNet + attention layers). In this vLLM release the prefix cache is off unless --enable-prefix-caching is passed — returning to a 120k-token conversation meant 131 s of re-reading. With the flag vLLM switches to mamba_cache_mode 'align' (marked experimental) and the return takes 0.5–1.2 s.
--gpu-memory-utilizationvLLM's memory budget covers the weights and the KV cache; XPU graphs (2.84 GiB) come on top. Usage also grows at runtime as the KV pool fills and on the first image. At 0.90, after a long context the card had 31.78 of 31.89 GiB in use. Hence a lower budget than on NVIDIA (0.94 there) — and the reason is more serious than performance:
When the B70 runs out of VRAM, the xe driver (TTM) evicts GPU buffers into host RAM instead of failing the allocation. That is kernel memory — the container's RAM limit does not cover it. Here 16 GB of host RAM vanished within a minute: [TTM] Buffer eviction failed, global OOM, no SSH for 23 minutes. Test anything that adds VRAM from a safe value upwards, keep ≥ 1 GiB headroom after the first image (the vision buffer is only allocated then, +0.5 GiB), and run a watchdog that kills the server on the first TTM line in dmesg.
MTP speculation (qwen3_5_mtp, 2 tokens) gave a single user ~51 tok/s decode. But it adds ~0.8 GiB of weights and grows the graphs from 2.8 to 4.75 GiB — at a 0.90 budget the card was full right after start (31.84 GiB). At three users one run hung for 31 s; at eight VRAM overflowed. On a multi-user server — no MTP. A ≤ 5-user variant with a lower budget will be tested separately.
With eight users each pasting ~50 pages per round the KV pool (191k tokens) overflows and vLLM preempts requests, recomputing them later — median time to first token 92 s, single ones up to 4.5 minutes. Zero errors, but you feel it under that load. The Z820 with a 431k pool never preempted in the same test. The per-conversation context (131,072) does not depend on the budget — the budget decides how many long conversations fit at once.
A llama.cpp SYCL build is sensitive to the oneAPI and runtime versions. A casual upgrade can break a working stack. We hold package versions and upgrade deliberately.
--parallel splits the contextIn llama.cpp -c is the total context, split across slots. -c 393216
--parallel 3 gives 3 × 131,072, not 3 × 393,216. KV in q8_0: 262k of
context ≈ 11.5 GiB on top of the weights.
-b 4096 -ub 1024 instead of the default 2048/512: prompt processing 579 → 716 tok/s (+24 %), time to first token −19 %. Decode unchanged — as expected.
--spec-type ngram-cacheThe server starts in 150 s and dies on the first request. Six attempts, always "connection refused", empty journal. The other n-gram variants are still to be tested.
237 W under load against a 250 W limit. Raising it to 275 W: +0.2 % for one user, +4 % for three. For 20 W and a louder fan — not worth it.
Recent llama.cpp keeps saved conversations in host RAM — --cache-ram, 8192 MiB by default. The state of one ~120k-token conversation with a q8_0 KV cache is 5.9 GB here. A container with 8 GB of RAM died on it after 13 minutes of a stress test (OOM killer), and with a smaller context it hung in swap for 20 minutes. Fix: keep --cache-ram below the container's memory limit (2048 here). After the change: 30 minutes, two users, conversations up to 122k tokens — no restart, peak RAM 5.3 GB.
The Qwen3.8-27B GGUF carries the MTP layer (blk.64.nextn.*) and llama.cpp has --spec-type draft-mtp. One user: 20.1 → 26.9 tok/s. Three at once: 26.7 → 14.3 tok/s aggregate — the draft is computed per slot. Plus 2.6 GiB of VRAM. On a server for several people — no.
Clean measurement, one user, nothing else on the server: at 15k / 60k / 120k tokens of context decode runs at 17.8 / 12.0 / 8.4 tok/s, prompt processing at 937 / 814 / 693 tok/s. Returning to the same conversation from cache — 1.5–2.6 s to first token. The fast Battlemage flash-attention path (llama.cpp PR #25222) only applies with an f16 KV cache and a single sequence: f16 KV gave us +16 % decode at 60k, but costs half the context.
We counted tokens in reasoning_content; vLLM 0.20.1 sends reasoning. Zero tokens seen, TTFT equal to the whole response, "6.66 tok/s" instead of the real 8.4. Always print how many stream chunks carry content. Zero = you are measuring noise.
max() of runs flips the rankingSpread reached 18 % (45.5 and 54.5 tok/s, same config). Best run instead of mean put the weaker machine ahead of the stronger one. Always the mean, always with the spread.
nvidia-smi reads the wrong cardOn a box with a B70 and an RTX 3060 the script reported 15 W and 210 MHz — the idle 3060. Intel telemetry lives in /sys/class/hwmon/*/ with name = xe: energy1_input (µJ), power1_cap, temp*_input.
memory.used to dropThe script compared the service's restart counter before and after the test. The server hung in swap for 20 minutes but never restarted — so it “survived the whole test”, while the last two rounds came back empty (0 tokens, 1,187 s to first token) instead of failing. Check tokens per reply, time to first token and cgroup OOM events, not only restarts.
When one user pastes 15k tokens, that prompt is processed in the same steps as the other user's decode — who then gets a token every few seconds. Hence “4 tok/s” with two users on the B70 and 2.4–3.2 tok/s with eight on the Z820. It is not a long-context slowdown — that needs a separate single-user measurement.
Average power from the counter difference at the start and end of a 30-minute test came out at 93 W, while sampling every second during the test showed 242 W. Read the counter often and sum the increments.
| Format / repo | Verdict for RTX 3090 (Ampere) and Arc B70 |
|---|---|
NVFP4 (unsloth/…-NVFP4 etc.) | Blackwell only (RTX 50xx) — out |
FP8 (Qwen/…-FP8) | Ada or Hopper — out on Ampere |
| NVFP4-MTP-GGUF | has an MTP head, but NVFP4 — useless on Intel |
| MLX | Apple Silicon format |
| int4 AutoRound / AWQ-INT4 | works on 3090 (Marlin path) — what we use |
GGUF UD-Q4_K_M / ggml-org | works on B70 via llama.cpp SYCL (our choice until 09-26) |
GPTQ int4 + MTP (SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16) | runs on the B70 via vLLM XPU (intel/llm-scaler-vllm), vision included — what we use since 09-26 |
On the 3090 no new format will help. On the B70 it was not the format but the engine — the same model in GPTQ int4 under vLLM XPU is 1.6–2.9× faster than GGUF under llama.cpp. Before pulling 20 GB, check your architecture supports it at all.