Hello, I'm
Sylar.
I build things, break things, and write about what I learn along the way.
投机解码:从拒绝采样到 Omni 模型
22 July 2026 2 min投机解码(speculative decoding)常被介绍成“用小模型猜、大模型验“。这个说法...
FP8 Attention on B300: Videos, Accuracy, Latency
3 July 2026 2 minAll results: Cosmos3-Nano on a single NVIDIA B300 (sm_103) in vLLM-Omni, official generation configs...
Full-Duplex in vLLM-Omni: A Map of the Design Space
30 June 2026 29 minThis is a map, not a verdict. Full-duplex interaction serving in vLLM-Omni is an active design — R...
All Posts
投机解码:从拒绝采样到 Omni 模型
22 July 2026 • 2 minute read投机解码(speculative decoding)常被介绍成“用小模型猜、大模型验“。这个说法...
FP8 Attention on B300: Videos, Accuracy, Latency
3 July 2026 • 2 minute readAll results: Cosmos3-Nano on a single NVIDIA B300 (sm_103) in vLLM-Omni, official generation configs...
Full-Duplex in vLLM-Omni: A Map of the Design Space
30 June 2026 • 29 minute readThis is a map, not a verdict. Full-duplex interaction serving in vLLM-Omni is an active design — R...
vLLM-Omni 量化推理实践
5 June 2026 • 3 minute read训练后量化是在不重训的前提下降低大型扩散 Transformer 显存与延迟成本的主...
That's all the posts so far!
Contact
You can find me on any of the following platforms: