Tendril
loading…
思考模式
temperature
max tokens
新对话
你自己的推理引擎
手写的 Qwen3 前向、paged KV cache、continuous batching、prefix caching
每条回答下方显示首字延迟、生成速度和缓存命中