ABOUT ME

-

Today
-
Yesterday
-
Total
-
  • DFlash 2 M4 Max 디코딩 실측 — 2.02배, Greedy 5/8 불일치
    AI Agent 2026. 8. 24. 19:38
    728x90
    반응형

    같은 M4 Max에서 별도로 측정한 DFlash 1의 1.83배와 DFlash 2의 약 2.02배를 조건과 함께 보여 주며, DFlash 2는 greedy token 출력이 8개 중 5개 달라 FAIL임을 표시한 도식입니다.
    DFlash 1은 suite post-prefill tok/s 기준 1.83배, DFlash 2는 prefill 제외 generation_tps 기준 약 2.02배였습니다. target·runtime·block·prompt·지표가 달라 우열 비교는 아니며, DFlash 2는 greedy token sequence가 5/8 불일치해 FAIL입니다.

    DFlash 2를 Apple M4 Max에서 돌렸습니다. 공식 CLI가 남긴 prefill 제외 generation_tps는 28.00에서 56.56으로 올랐고, 표시 로그 5개의 speedup 기하평균은 약 2.02배였습니다. 그런데 temperature 0 Greedy 검사에서는 8개 prompt 중 5개의 token sequence가 달랐습니다.

    먼저 결과

    판정 축 DFlash 2 M4 Max 결과
    디코딩 관측 prefill 제외 generation_tps 28.00 → 56.56, 약 2.02×
    사전등록 end-to-end 지표 NOT_CAPTURED_HOLD
    Greedy token sequence 3/8 일치, 5/8 불일치
    종합 판정 FAIL

    재현 코드와 prompt SHA-256은 DFlash 2 공개 runner에, full revision과 판정 순서는 DFlash 2 protocol에 올렸습니다.

    무엇을 확인했나

    Inco AI의 DFlash 2 발표는 NVIDIA H200과 SGLang에서 end-to-end throughput 2.67–3.43배를 보고합니다. 제 조건은 M4 Max와 MLX였기 때문에 그 숫자를 그대로 재현하는 실험으로 잡지 않았습니다. 같은 target과 prompt 안에서 baseline 대비 디코딩 속도가 오르는지, Greedy output이 같은지를 따로 확인했습니다.

    DFlash 2의 draft model은 block 안의 각 위치에 여러 token 후보를 병렬로 만듭니다. 후보 경로 선택기(candidate path selector)가 앞뒤가 이어지는 경로를 고르고, 두 탭 동적 깊이별 합성곱(two-tap dynamic depthwise convolution)이 현재 위치와 바로 앞 위치의 정보를 섞습니다. 마지막 확정은 target model이 맡습니다.

    이 설명은 공식 구조에 대한 요약입니다. 이번 실행에서는 path selector나 convolution을 하나씩 끄는 ablation을 하지 않았습니다.

    어떻게 측정했나

    항목 조건
    하드웨어 Apple M4 Max, 64GB 통합 메모리, 40-core GPU
    운영체제 macOS 26.6.2, build 25G83
    런타임 Python 3.14.6, dflash==0.1.0, mlx==0.32.0, mlx-lm==0.31.3
    target / draft Qwen3.8-27B target 4bit / DFlash 2 draft runtime 4bit
    prompt 고정 revision의 GSM8K test 8문항, selector seed 42
    성능 측정 max_new_tokens=256, temperature 1, top-p 0.95, top-k 20, reasoning xhigh, block 5
    parity 측정 같은 8문항, temperature 0, reasoning xhigh, 전체 token sequence 비교
    반복 8문항 full benchmark를 순차 5회, 반복 사이 cooldown 60초

    공식 MLX CLI는 각 prompt에서 baseline을 먼저 실행하고 DFlash 2를 이어서 실행합니다. 각 full benchmark에는 CLI 내부 warm-up이 들어 있습니다. 다섯 번은 같은 장비에서 순서대로 돌렸으며 통계적 독립성은 확인하지 않았습니다.

    첫 performance run은 숫자를 하나도 남기기 전에 HFValidationError로 멈췄습니다. local draft snapshot 경로를 repo ID로 해석한 것이 원인이었습니다. 실패 로그는 보존했고, exact SHA cache를 확인한 뒤 offline repo ID 경로로 바꿨습니다. model revision, prompt, sampling, block size와 반복 수는 그대로 뒀습니다.

    사전등록한 1차 지표는 end-to-end wall time이었습니다. 실제 raw에 남은 값은 prefill 제외 generation_tps였습니다. 그래서 속도 축은 OBSERVATION_ONLY로 남겼고, Greedy mismatch가 종합 FAIL을 결정했습니다.

    결과

    반복 AR generation_tps DFlash 2 generation_tps 표시 speedup
    1 28.01 55.78 1.99×
    2 27.12 57.06 2.10×
    3 28.24 56.41 2.00×
    4 28.18 56.51 2.00×
    5 28.47 57.03 2.00×

    AR은 28.00 ± 0.52, DFlash 2는 56.56 ± 0.53 generation_tps였습니다. CLI가 두 자리로 남긴 speedup 5개의 기하평균은 약 2.02배였습니다.

    Greedy token sequence는 3/8이 일치하고 5/8이 불일치해 사전등록 기준상 FAIL이었습니다. 불일치한 prompt index는 0, 1, 5, 6, 7입니다. 후속 진단은 이 다섯 prompt만 같은 process에서 반복했습니다. 각 경로 내부 결과는 같았고 두 경로 사이의 mismatch 5건도 다시 나타났습니다. fresh process와 나머지 세 prompt는 이 진단에 포함하지 않았습니다.

    DFlash 1 숫자와 직접 나눌 수 없는 이유

    별도 프로토콜 참고 DFlash 2 DFlash 1
    target Qwen3.8-27B 4bit Qwen3.5-27B 4bit
    baseline → DFlash 28.00 → 56.56 generation_tps 28.33 → 52.84 tok/s
    내부 speedup 약 2.02× 1.83×
    block / prompt block 5 / GSM8K 8개 block 16 / GSM8K 10개
    correctness 5/8 불일치 → FAIL Greedy parity 미검증

    DFlash 2의 약 2.02배와 DFlash 1의 1.83배는 각각 자기 baseline과 비교한 값입니다. Qwen3.8과 Qwen3.5, dflash==0.1.0dflash-mlx==0.1.8, block 5와 16, 8문항과 10문항이 다릅니다. speedup 집계도 표시 로그 기하평균과 suite median으로 갈립니다.

    정확성 상태도 다릅니다. DFlash 2는 5/8 mismatch를 측정했고 DFlash 1은 token ID나 hash를 남기지 않았습니다. DFlash 1의 미검증 상태를 PASS로 채울 수 없습니다. DFlash 2의 전체 target+draft memory도 확보하지 못해 메모리 우열은 비교하지 않았습니다.

    남은 것

    이번 결과는 M4 Max 한 대, Qwen3.8-27B 한 조합, GSM8K 8문항, 순차 n=5에 한정됩니다. 다른 block size·precision·task·concurrency와 sampling distribution은 확인하지 않았습니다. pmset -g therm에는 warning이 없었지만 실제 온도를 잰 것은 아니어서 thermal 영향은 UNKNOWN입니다.

    5/8 mismatch의 원인도 아직 모릅니다. 다음 진단에서는 첫 불일치 위치에서 one-token target logits와 block verify logits를 맞대고, Gated DeltaNet state가 rollback 전후에 같은지 확인할 예정입니다.

     

    github: https://github.com/joonyeonglim/dflash-m4-max-benchmarks

     

    GitHub - joonyeonglim/dflash-m4-max-benchmarks: DFlash 1/2 benchmarks on Apple M4 Max with methods, sanitized results, and DFlas

    DFlash 1/2 benchmarks on Apple M4 Max with methods, sanitized results, and DFlash 2 parity analysis. - joonyeonglim/dflash-m4-max-benchmarks

    github.com

     

    728x90
    반응형
Designed by Tistory.