|
|
|
I'm currently a senior software engineer working at AMD ROCm (previously at Huawei Ascend), building vLLM inference engine for GPU/NPU software ecosystem (focusing on multi-modality inference / structured output / OOT hardware extensibility).
|
|
|
I'm currently a senior software engineer working at AMD ROCm (previously at Huawei Ascend), building vLLM inference engine for GPU/NPU software ecosystem (focusing on multi-modality inference / structured output / OOT hardware extensibility).
A high-throughput and memory-efficient inference and serving engine for LLMs
Community maintained hardware plugin for vLLM on Ascend
This repo is used for archiving my notes, codes and materials of cs learning.
A curated collection of Claude Code agent skills that accelerate the entire vLLM development lifecycle.
Forked from SemiAnalysisAI/InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究…
Python