Loading...
Loading...
Compare original and translation side by side
Min LatencyBalancedMax Throughputtrtllm-serve --configdocs/source/features/examples/llm-api/Min LatencyBalancedMax Throughputtrtllm-serve --configdocs/source/features/examples/llm-api/speculative_configdecoding_type: MTPexamples/configs/speculative_configdatabase.pyMin LatencyBalancedMax ThroughputLow LatencyHigh Throughputspeculative_configexamples/configs/decoding_type: MTPspeculative_configdatabase.pyMin LatencyBalancedMax ThroughputLow LatencyHigh ThroughputConfigSourceLaunch commandConfigSource used as starting pointWhat to benchmark配置来源启动命令配置用作起点的来源基准测试内容Min LatencyBalancedMax ThroughputMin LatencyBalancedMax Throughputexamples/configs/database/lookup.yaml(model, gpu, isl, osl, concurrency, num_gpus)database.pyexamples/configs/database/lookup.yaml(model, gpu, isl, osl, concurrency, num_gpus)database.pyexamples/configs/curated/lookup.yamldatabase/curated/qwen3-disagg-prefill.yaml*-latency.yaml*-throughput.yamlexamples/configs/curated/lookup.yamldatabase/curated/qwen3-disagg-prefill.yaml*-latency.yaml*-throughput.yamldocs/source/deployment-guide/examples/models/core/docs/source/legacy/examples/models/core/deepseek_v3/README.mddocs/source/deployment-guide/examples/models/core/docs/source/legacy/examples/models/core/deepseek_v3/README.mdmax_batch_sizemax_num_tokensmax_seq_lenenable_attention_dpattention_dp_config.*kv_cache_config.free_gpu_memory_fractionmoe_expert_parallel_sizemoe_config.backendstream_intervalnum_postprocess_workerscuda_graph_config.max_batch_sizebatch_sizesreferences/knob-heuristics.mdmax_batch_sizemax_num_tokensmax_seq_lenenable_attention_dpattention_dp_config.*kv_cache_config.free_gpu_memory_fractionmoe_expert_parallel_sizemoe_config.backendstream_intervalnum_postprocess_workerscuda_graph_config.max_batch_sizebatch_sizesreferences/knob-heuristics.mdtrust_remote_code: truemax_num_tokenstrust_remote_code: truemax_num_tokens