Perplexity announced today that Portable Computer — the fully local version of its Computer agent — is now available on Windows PCs with NVIDIA GeForce RTX or RTX PRO GPUs, distributed through the Microsoft Store. NVIDIA's own account amplified it within the hour: "Powerful AI agents are becoming easier to run right on your PC. Together with @Perplexity_AI, we're making local AI more private, accessible and useful on @Windows NVIDIA RTX PCs." Both companies' framing is identical and specific: run the harness, the agents, and the models on-device, work with local files and connected apps without sending tasks to the cloud, and escalate to frontier cloud models only when a task needs them.
Two similarly named products that are not the same thing
Worth clearing up before anything else, because Perplexity's own naming invites the confusion: Portable Computer is not Personal Computer. Personal Computer is the older product — Perplexity's cloud-coordinated Mac agent, in Max-subscriber beta from mid-April 2026 and opened to all Mac users on May 7, then given a hybrid mode on September 1 that runs sensitive steps on Apple Silicon locally while the rest goes to the cloud. Portable Computer is a separate, newer product announced in late August: not hybrid-by-default but fully local by design, with the entire stack — orchestrator model, subagent model, and agent harness — running on the user's own machine, and cloud invoked only as an explicit, permissioned escalation rather than a default coordination layer. macOS is not on Portable Computer's roadmap at all; it launched for NVIDIA hardware specifically, which is the part that makes today's news a continuation of an existing partnership rather than a new one. Two products, from the same company, three weeks apart in this post's timeline alone, sharing four of six words in their names and pointed at different hardware — the kind of overlap that makes it worth checking which one a headline actually means before citing either.
Three platforms in three weeks
Portable Computer's own rollout has moved fast enough that today's Windows release is really the third leg of a short sprint. It shipped first on August 25 for NVIDIA's DGX Spark, the 4,679 AI desktop to any sufficiently equipped consumer GPU) while holding the VRAM requirement constant.
The model story is the part this blog has already been tracking
The more interesting detail is which models Portable Computer actually runs, because it's a live test of an architecture this blog flagged a month ago. At launch, users choose between Perplexity's own PPLX 27B — a model Perplexity post-trained specifically for its own harness — and Qwen 3.8 27B, Alibaba's open-weight model tuned for RTX hardware. Perplexity has also said NVIDIA Nemotron 3.5 Lightning support is coming soon.
That's a direct continuation of a story this blog covered when NVIDIA shipped Nemotron 3.5 Lightning on August 11: a 30B mixture-of-experts model with roughly 3B active parameters, open-weighted, explicitly not pitched as a frontier generalist but as "the specialized executor inside" someone else's larger agent system, shipped alongside NeMo Switchyard, an open routing library for deciding when to call it. That post's read on NVIDIA's business logic was that giving away a capable small executor model costs NVIDIA nothing if the alternative is the same inference happening on someone else's metered API — "NVIDIA doesn't need Nemotron to win. It needs the category of locally-run specialist executors to win." A month later, one of the more visible consumer AI products on the market is building exactly that category: a local orchestrator with a swappable small-model slot, currently filled by Perplexity's own model or Qwen, with NVIDIA's own entry slated to join the roster it's competing inside. This blog also flagged the same pattern four days ago in Sakana's Fugu Ultra v2 launch, which leans on Nemotron models as cheap pool workers while explicitly excluding frontier flagships from its own orchestration pool. Two unrelated companies, in the same two weeks, both building the "cheap local executor, NVIDIA fills the slot" architecture that a NVIDIA product launch predicted for them — which is either NVIDIA reading the market correctly or NVIDIA's own giveaway strategy manufacturing the market it predicted, and there's no way from the outside to tell those apart.
It's also worth naming who wins from PPLX 27B specifically: Perplexity, not NVIDIA. Offering its own post-trained model as the default choice, with Qwen and eventually Nemotron as the swappable alternatives, keeps Perplexity in control of the harness's default behavior even on hardware NVIDIA supplies — the model layer stays contestable, but the harness that decides which model runs stays Perplexity's.
An internal benchmark, on an internal test
Perplexity's own figures for the model choice: on its 53-task "Local Knowledge Work Bench" — internal, covering tasks like deep research, financial analysis, and document creation — Qwen 3.8 27B scores 82.6%, and Perplexity's own PPLX 27B scores 85.4%. Both numbers come from Perplexity's own benchmark, scored by Perplexity, on tasks Perplexity wrote, comparing Perplexity's model favorably against the alternative it also offers. None of that makes the underlying result false — a company post-training a small model specifically for its own harness plausibly should outperform a general-purpose model of the same size on tasks drawn from that harness's own use cases. But it's the same structure this blog flagged in Desert Ant Labs' launch nine days ago and in Sakana's Chartography claim four days ago: a single-source number, on a benchmark the publisher also designed, is a claim worth holding as provisional rather than treating a 2.8-point gap as established fact until someone outside Perplexity runs the same 53 tasks.
What "zero token cost" costs
Every piece of coverage of today's launch, including Perplexity's and NVIDIA's own, leads with the same framing: local steps run at zero marginal token cost. That's true as stated and worth taking seriously — this blog made exactly this argument in August about Liquid AI's LFM2.5-2.6B, a 2.6B model that runs an agent on a phone people already own, where removing the meter genuinely changes what you'd build rather than just what you'd pay. Portable Computer's version of "zero token cost" sits on a meaningfully different foundation. Reaching it requires a Perplexity Pro subscription at 200 a month, plus hardware: either NVIDIA's 1,500-to-$2,000-plus range even before the rest of the machine around it. A phone someone already owns and a 24GB workstation GPU they'd need to specifically buy are not the same kind of "free." For a light user, the per-token cost of a cloud API is very likely cheaper than the hardware outlay required to eliminate it — "zero token cost" describes the marginal cost of each additional task once the fixed cost is paid, not the total cost of using Portable Computer at all, and the coverage racing to repeat the "zero" framing mostly isn't making that distinction.
The one design choice in the launch that reads as more than marketing is the escalation mechanic itself: every step starts on-device, and sending anything to one of the 15-plus cloud models requires a separate, explicit approval after a privacy check, with that approval covering only the one transfer rather than persisting across steps or sessions. That's a real, specific commitment — friction a company chasing usage metrics would normally want to minimize, not add — and it's a genuinely different design than a harness that silently decides when to go to the cloud on the user's behalf.
Two companies, two different things to sell
NVIDIA and Perplexity aren't really competing for the same dollar here, which is what makes the partnership durable rather than opportunistic. NVIDIA sells the GPU that makes local inference possible at all — every Portable Computer install on RTX hardware is a sale NVIDIA already made, regardless of which model runs on it, which is the same logic this blog traced through Nemotron 3.5 Lightning's own giveaway economics. It's also not the first time these two companies have been in the same room this quarter: Perplexity was already a founding member of NVIDIA's Open Secure AI Alliance, announced in late July, a coalition built on the same underlying bet — that value accrues to whoever controls the infrastructure open models run on, not necessarily to whoever trains the model. Perplexity sells the subscription and, longer-term, control of the harness layer that decides which model — its own, an open one, or eventually NVIDIA's — actually runs. Both companies get to point at the other's brand while collecting a different bill. That's not a criticism particular to this launch; it's just worth stating plainly, the same way this blog has stated it for every company that's argued this month that its own product happens to be the correct answer to a debate it also has a stake in. What makes today notable isn't the Windows port itself — that was the predictable next stop on a three-week rollout — it's that the "cheap, swappable, locally-run executor" architecture NVIDIA was still selling as a thesis in an August product launch now has a live, shipping example built by someone else, with NVIDIA's own model queued up to join it rather than compete against it.