News: 0185381344

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Perplexity Will Open Source Its Faster Lily AI Engine For Apple Silicon

(Thursday September 03, 2026 @11:00AM (BeauHD) from the optimized-inference dept.)


[1]BrianFagioli writes:

> Perplexity has built a local artificial intelligence engine [2]designed specifically for Apple silicon and the Qwen3.6-35B-A3B model . Called Lily, the engine uses a Rust runtime and custom Metal kernels, with neither PyTorch nor MLX in its execution path. Perplexity [3]says Lily averaged 23 percent faster prompt processing and 35 percent faster token generation than MLX-LM on an M5 Max MacBook Pro with 128GB of unified memory. Lily is more specialized than MLX-LM, which supports a much wider range of models and architectures. Perplexity says it plans to release Lily as open source, but the code is not available yet, leaving its performance claims dependent on internal testing for now.



[1] https://slashdot.org/~BrianFagioli

[2] https://www.perplexity.ai/hub/blog/optimizing-on-device-inference-for-apple-silicon

[3] https://www.perplexity.ai/hub/blog/optimizing-on-device-inference-for-apple-silicon



M5 Max MacBook Pro with 128GB (Score:2)

by thesjaakspoiler ( 4782965 )

Now that is promising for many people.

Re: M5 Max MacBook Pro with 128GB (Score:2)

by beelsebob ( 529313 )

I mean, if youâ(TM)re running AIs, youâ(TM)re gonna want 128GB at least, and a fast GPU. At that point, it being a MacBook Pro isnâ(TM)t going to make much difference. Itâ(TM)s gonna cost $7000, which I grant you is a lot of money, however⦠now find me another laptop with a GPU around the performance of the laptop 5080, and 128GB of RAM the GPU can access.

Re: (Score:2)

by dfghjk ( 711126 )

"...now find me another laptop with a GPU around the performance of the laptop 5080, and 128GB of RAM the GPU can access."

For the market that must run local AI as fast as possible on a portable device. Wonder what application needs that?

Re: (Score:2)

by beelsebob ( 529313 )

I mean, if you like, we can loosen the requirements and you can look for a desktop. Then the mac costs $5099. I'm not sure you're going to get a GPU with 128GB for that much, let alone the whole system.

Re: (Score:2)

by tlhIngan ( 30335 )

> I mean, if youÃ(TM)re running AIs, youÃ(TM)re gonna want 128GB at least, and a fast GPU. At that point, it being a MacBook Pro isnÃ(TM)t going to make much difference. ItÃ(TM)s gonna cost $7000, which I grant you is a lot of money, howeveræ now find me another laptop with a GPU around the performance of the laptop 5080, and 128GB of RAM the GPU can access.

A laptop or desktop with 128GB of RAM will likely be around $7000, with RAM costs being what they are. Even a 64GB la

but promising what? (Score:2)

by dfghjk ( 711126 )

What is it promising?

Optimise all the things! (Score:2)

by nickovs ( 115935 )

They built a custom engine for a specific model and it went 30% faster. That's great, but it's inflexible and of limited use.

Of course these days, coding agents make designing and writing new code very cheap. With the right tools and toolchains it shouldn't be hard to automate the process of building a custom engine for each new model that comes along. The cost of doing this would be small compared to the cost of training the model in the first place, so really we should just be doing this as a matter of co

Re: Optimise all the things! (Score:3)

by beelsebob ( 529313 )

I somewhat disagree. They developed an engine using somewhat different building blocks. The first iteration of their design works on only one model, but guess what - a lot of the blocks will transfer to other AIs. Iâ(TM)d bet that this can be expanded to run at least other versions of Qwen, and plausibly other LLMs.

Ultimately, all software is about trade offs, and here theyâ(TM)ve made a somewhat different trade from others, but given how popular high end Apple silicon devices are becoming for

Re: (Score:2)

by dfghjk ( 711126 )

"That's great, but it's inflexible and of limited use."

It won't be able to kill people nearly fast enough.

"...coding agents make designing and writing new code very cheap. "

Writing new code has always been "very cheap" as long as you don't care if it works.

"With the right tools and toolchains it shouldn't be hard to automate the process of building a custom engine for each new model that comes along."

Sure Sam, it's all easy. That's why no one has to work any more.

"The cost of doing this would be small comp

love, n.:
When you don't want someone too close--because you're very sensitive
to pleasure.