Databricks Proteus writes GPU kernels 1.8x to 5.2x faster than vLLM
Databricks says its Proteus agent harness generated Qwen 3.5 122B kernels 1.8x to 5.2x faster than vLLM on NVIDIA B200 GPUs, and that validation, not generation, was the hard part.
Read More