PagedAttention on 安橙的博客

PagedAttention on 安橙的博客https://blog.ans20xx.com/tags/pagedattention/Recent content in PagedAttention on 安橙的博客Hugo -- 0.163.3zhSat, 20 Jun 2026 00:00:00 +0000Day 31 · PagedAttention & vLLMhttps://blog.ans20xx.com/posts/ai/day31/Sat, 20 Jun 2026 00:00:00 +0000https://blog.ans20xx.com/posts/ai/day31/学习 PagedAttention 与 vLLM 的核心机制:为什么 KV Cache 会浪费显存,如何用 block table 管理逻辑块到物理块的映射,copy-on-write 如何支撑并行采样和 beam search,以及这些机制如何服务高吞吐 LLM serving。Day 32 · vLLM 实战https://blog.ans20xx.com/posts/ai/day32/Sat, 20 Jun 2026 00:00:00 +0800https://blog.ans20xx.com/posts/ai/day32/动手部署一个 7B 模型到 vLLM,开启 OpenAI 兼容 API,学习 --max-num-seqs 与 --gpu-memory-utilization 的调参方法,并建立推理服务压测与排错流程。