Today we’re open-sourcing Lily, the local inference engine we built for hybrid compute in Perplexity Computer. Lily is specialized for Qwen3.6-35B-A3B on Apple silicon, built so on-device compute doesn’t bottleneck Computer tasks. Read more: ↧ Optimizing On-Device Inference for Apple Silicon A custom local engine that improves prefill and decode throughput