Companies & Products

Perplexity launches Portable Computer, a local agent with no token fees

Perplexity released Portable Computer on Tuesday, an agentic tool that keeps its models, your files and your work on your own machine instead of in the cloud. It’s the on-device version of the cloud agent the company shipped in February, and ZDNet reports it’s restricted to paid Pro, Max, Enterprise Pro and Enterprise Max accounts. The first release runs on Linux only.

The pitch is money. Work done on the device doesn’t consume billing credits or tokens, and only the parts of a task that hop online eat into your allowance. That’s the same arithmetic behind running your own models rather than renting an API, applied to an agent harness.

Sensitive data therefore never leaves the device without permission, and local models carry no inference fee: the system is private and cost-effective by construction.

Perplexity, technical blog post, quoted by ZDNet

The catch is the hardware bill. On Linux the agent needs an Nvidia DGX Spark or another Linux box with an Nvidia RTX GPU, running DGX OS or Ubuntu on ARM or x64. SiliconANGLE reports the DGX Spark pairs a Blackwell-architecture graphics card with a 20-core processor and 128GB of memory. Windows support is due in September, and there the floor is an RTX card with at least 24GB of VRAM, which VentureBeat puts at roughly a GeForce RTX 3090 or newer.

Two local models are available at launch. One is Qwen 3.8 27B, the open-source release; the other is PPLX 27B, a post-trained version Perplexity tuned for the DGX Spark with a multi-token prediction mechanism meant to speed up prompt processing. Nvidia’s Nemotron 3.5 Lightning is queued next, and Nvidia describes it as a 30-billion-parameter mixture-of-experts model that delivers up to 4x faster output speed than others in its class.

Perplexity’s own benchmark numbers, reported by VentureBeat, put the local setup against two free agent harnesses, Pi and Hermes.

BenchmarkPerplexity on DGX SparkPiHermes
Local knowledge work, Qwen 3.8 27B82.6%77.6%74.0%
Local knowledge work, PPLX 27B85.4%n/an/a
BrowseComp web research66.7%50.2%43.9%
Multimodal document understanding65.1%13.9%34.6%
Perplexity’s internal results, as reported by VentureBeat, 25 August 2026.

Nobody outside the company has reproduced any of that, so treat it as vendor testing. The escalation numbers are the more honest ones. VentureBeat reports the fully local Qwen model scored 59.6% on a coding benchmark at essentially zero marginal cost, and that letting it call out to a Claude Opus 5 advisor lifted the score to 73.0% at an estimated $0.415 per task.

Context length is the other constraint. SiliconANGLE reports that Perplexity found Qwen 3.8 27B struggles with prompts beyond 100,000 tokens despite advertising a far larger window, so the agent ships with a compaction feature that summarises long prompts back under that mark. The guardrails are more conventional: a sandbox that walls the model off from irrelevant parts of your operating system, blocked network connections it didn’t ask for, and a permission prompt before anything goes to an outside tool.

What’s harder to read is the Nvidia framing. The Register’s Thomas Claburn tied the launch’s heavy Nvidia emphasis to reports that the chipmaker is weighing an investment valuing Perplexity above $30 billion. SiliconANGLE adds that Nvidia has reportedly floated licensing Perplexity’s technology and hiring key employees. Claburn’s sharper question is how Perplexity charges for a local harness when the ones it benchmarked against cost nothing.

September answers two of these. That’s when Windows machines with RTX cards get access, and when the subscriber base widens past people who own a DGX Spark or a 24GB GPU.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *