News: 0185154330

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Perplexity and Nvidia Launch Fully Local AI Agent With Zero Token Costs

(Tuesday August 25, 2026 @05:01PM (BeauHD) from the cloud-is-optional dept.)


Perplexity and Nvidia have launched " [1]Portable Computer ," a local-first version of Perplexity's agent platform that [2]runs AI models, files, tools, and workflows directly on Nvidia-powered Linux hardware . Local tasks incur no token charges and keep data on-device by default, with users asked for permission before the system escalates a step to a cloud model. VentureBeat reports:

> For Nvidia, which has spent the past two years selling the world on trillion-dollar AI data centers, the announcement signals something subtler but strategically important: the chipmaker believes local AI has crossed a threshold from hobbyist curiosity to practical tool -- and it wants to sell the hardware that runs it.

>

> "Local AI reached an inflection point," said Nader, Nvidia's director of developer technology, who focuses on developer tooling and open source. "For the longest time, it was hobbyists and enthusiasts, and they were running these quantized models that were quantized down to be super tiny... And while that's cool, it's not super practical. But all that changed with a lot of these new open source models that have come out that are super useful."

>

> [...] Portable Computer arrives today for Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Windows support following in September. Any RTX GPU with at least 24GB of VRAM -- roughly a GeForce RTX 3090 or newer -- clears the bar, a threshold Nate called "sort of the floor where we really want to make sure that we can deliver a great experience, but balance that with making it broadly available."



[1] https://www.perplexity.ai/hub/products/portable-computer

[2] https://venturebeat.com/infrastructure/perplexity-partners-with-nvidia-to-launch-portable-computer-a-fully-local-ai-agent-with-zero-token-costs



Fully local my arse. (Score:2)

by devslash0 ( 4203435 )

Let me know when those models stop calling home with telemetry and can run at full capability without access to the internet.

Re:Fully local my arse. (Score:4, Insightful)

by bsolar ( 1176767 )

> Let me know when those models stop calling home with telemetry and can run at full capability without access to the internet.

Try something like Ollama + OpenCode. Or even Claude Code with CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC, if you trust them enough.

Re: Fully local my arse. (Score:2)

by devslash0 ( 4203435 )

I'm never going to trust a model-provided flag, mate. Full firewall, completely cut off from the internet is the only way. Sadly, many models won't work at all when isolated.

Re: (Score:2)

by EvilSS ( 557649 )

> Sadly, many models won't work at all when isolated.

Which ones? Name some.

Re: (Score:2)

by bsolar ( 1176767 )

> I'm never going to trust a model-provided flag, mate. Full firewall, completely cut off from the internet is the only way. Sadly, many models won't work at all when isolated.

Just ran a quick test and e.g. ollama launch claude --model qwen3-coder-next works offline without issue, as long as you already have the model downloaded locally.

do you need internet access to function?

No, I don't need internet access to function. I can work offline with:

- Reading/writing files on your local system

- Running local commands (git, npm, builds, tests, etc.)

- Using tools like LSP for code navigation

- Managing tasks and workflows

- Creating artifacts and editing notebooks

I do have

Re:Fully local my arse. (Score:4, Interesting)

by EvilSS ( 557649 )

> Let me know when those models stop calling home with telemetry and can run at full capability without access to the internet.

What local models are sending telemetry? And how exactly are they doing that? None of the ones I use do that, nor can they, and they work perfectly offline.

Re: (Score:2)

by abulafia ( 7826 )

This is just Qwen with an installer.

I mean, yeah, wouldn't surprise me if this thing phones home. But I wonder how many folks there are who own a 3090 or better who also can't handle Ollama or a firewall.

Re: Fully local my arse. (Score:2)

by devslash0 ( 4203435 )

Setting up the firewall isn't the problem here. The problem is that when you do set it up, those models often refuse to work at all.

Re: (Score:2)

by anomaly256 ( 1243020 )

It's not the models refusing, it's the harness. Open weight models work just fine without an internet connection. The same can't be said for these vendor-provided agent harnesses with telemetry baked in though

Re: (Score:2)

by Fly Swatter ( 30498 )

You've been told that AI can not run locally and you believed it. It can, just not with your available hardware.

For entertain purposed only, I've become addicted to Stable Diffusion for image generation. If you use the less capable old models it runs completely local on a 4GB GPU I bought over 6 years ago. Is it fast? No. Does it work in a self contained environment? yes.

Just for this hobby I would have bought a new high end GPU had the prices not rocketed up to stupidsville.

..."trillion-dollar AI data centers" (Score:1)

by Valgrus Thunderaxe ( 8769977 )

Citation needed for these trillion-dollar data centers (even just one).

Re: (Score:2)

by OrangeTide ( 124937 )

Collectively all the AI data centers represent about a trillion in capital expenditure. Rather than the plural form meaning there are multiple data centers each other over a trillion dollars.

(also, I think someone edited the trillion-dollar part out. I didn't see it in the summary. but I believe you that it was probably worded that way earlier today)

"Broadly available" is a stretch (Score:2)

by Powercntrl ( 458442 )

*looks at prices of that GeForce RTX 3090 with at least 24GB of VRAM*

Nah, I'm good fam. Not selling a kidney to run a local AI model.

Re: (Score:2)

by Locke2005 ( 849178 )

You're not the target audience. The companies getting tired of paying $50K for tokens are the target audience. Believe me, the hardware cost is trivial compared to how much the AI companies are raping businesses for.

Re: (Score:2)

by MobyDisk ( 75490 )

I suspect a company paying $50k/month for tokens is going to need a whole lotta RTX 3090s. I also suspect that this 24GB model is not as powerful or efficient as something they can get in the cloud. But with that said, this is basically what I do at home - run one locally for common tasks. I keep waiting for the local ones to achieve parity with the cloud versions, but RAM is so expensive because of the data centers, that I'm not sure it will happen before SkyNet launches the nukes.

Re: (Score:2)

by JaredOfEuropa ( 526365 )

Exactly. Especially as it seems that the AI companies are selling their tokens at below cost. Small-scale installations for local models may make sense for individuals who want to keep their data isolated, but they do not make economic sense for companies.

Re: (Score:2)

by spacepimp ( 664856 )

This seems to still require a Perplexity Pro, Max, Enterprise Pro, and Enterprise Max subscription for it to work.

Re: (Score:2)

by Tailhook ( 98486 )

An RTX 3090 24GB is about $2150 new.

I've spent more on gaming rigs a couple times: it's not some unfathomable amount of money.

Re: (Score:2)

by Fly Swatter ( 30498 )

This a prelude for the RTX Spark due to be released in the fall. It is also how AI should have been done from the beginning if, you know, 'local' hardware was capable at the time.

24GB VRAM ?! (Score:2)

by greytree ( 7124971 )

24GB VRAM ?

Hear that noise ? That's rich prople laughing at us.

Marginally Better Than Buzzword Bingo (Score:2)

by Grady Martin ( 4197307 )

> And while that's cool, it's not super practical. But all that changed with a lot of these new open source models that have come out that are super useful.

Dumb it down for me, nerd man. I have no idea what all this techno mumbo jumbo means.

Re: (Score:2)

by Fly Swatter ( 30498 )

AI inference is all math calculations, math numbers can be rounded off to fewer significant digits but you lose precision. The data within LLM models can be rounded down to less precise numbers, sort of like rounding the number 2.4368966734 to 2.44 requires less memory to contain. In this way big models can be made much smaller to fit within less memory but are less precise - the results become poor compared to the original big model.

This process is called quantizing.

3090 or newer? (Score:2)

by GrahamJ ( 241784 )

No not just newer, only xx90 cards have at least 24GB of VRAM and they all cost an arm and a leg. And you can already run models on them.

The man who follows the crowd will usually get no further than the crowd. The
man who walks alone is likely to find himself in places no one has ever been.
-- Alan Ashley-Pitt