UC Berkeley's K-search: Auto-Translating CUDA Kernel Optimizations to Apple's MLX
berkeley_ai · x · 2026-07-30
As the AI hardware ecosystem fragments, transferring mature kernel optimizations (like NVIDIA's CUDA) to new chips (like Apple Silicon) typically requires extensive manual rewrites.
To solve this, UC Berkeley and IBM Research introduced K-search, an evolutionary kernel search framework. It automatically translates decades of CUDA optimization knowledge into native strategies for Apple's MLX framework. Benchmarks show their translated attention code achieves 0.97x the speed of Apple's native implementation. The project is open-source.
More from Infra
- TensorSharp Adds Multi-GPU Tensor Parallelism, Boosting Local GGUF Inference Speeds — fuzhongkai · 2026-07-30
- NVIDIA Open-Sources PyCuTe: Pure Python Layout Algebra for CUTLASS — asdf1234_0 · 2026-07-30
- Advantech Edge Device Powered by Nvidia Thor Runs RealSense GMSL Cameras — chrismatthieu · 2026-07-30
- Together Offers Lowest Price and Highest Cache Hit Rate for Kimi K3 on OpenRouter — zhyncs42 · 2026-07-30
- Samsung's Q2 Operating Profit Surges 1,800% to Record High Amid AI Chip Boom — Polymarket · 2026-07-30
- The Race for Power: Assessing Global Electricity Production for AI — lemire · 2026-07-30