A startup called WiCi Technology is taking a different approach to client-side AI computing. Instead of putting a high-end GPU inside every PC or edge device, the company wants to make GPUs accessible wirelessly across a local network.
WiCi is preparing to launch WiCi One, a dedicated wireless GPU system that combines standard PC hardware with a software stack designed to make high-performance GPU computing accessible to laptops, smartphones and other devices without requiring a physical GPU connection. The company is effectively treating the wireless network as an extension of the computer’s I/O subsystem.
WiCi’s approach is to put the accelerator close to the user while separating it from the client device. And it supports gaming-oriented GPUs, not the more expensive AI accelerators from NVIDIA. Of course, cheap is a relative term. An RTX 5090 costs $5,000, if you can get one. An RTX 6000 enterprise accelerator costs about $12,000.
Still, it puts generative AI workloads on the client, where they typically have not been placed. Processing a large language model or other generative AI is typically a server-side process. That means processing and moving data over the network. If it is possible to process data on the client where it would be consumed, that saves on network traffic and server workload.
The initial WiCi One configuration combines an Intel Core Ultra 7 255H processor with an NVIDIA GeForce RTX 5060 Ti and 16GB of GPU memory ($499). A higher-performance configuration is planned around NVIDIA’s RTX 5090 with 32GB of memory. Both systems are designed around Wi-Fi 7, including 4×4 MIMO and 320MHz channels, with NVMe storage providing additional capacity for AI workloads.
The wireless connection is the key architectural challenge. Conventional GPU computing relies on extremely high-bandwidth, low-latency interfaces such as PCI Express. Simply moving GPU operations across a wireless network would introduce latency and bandwidth bottlenecks that could undermine performance.
WiCi has developed a software layer called the WiCi Protocol, which incorporates caching, deduplication, scheduling and pipelining techniques intended to minimize the amount of data that has to run over the wireless connection.
Its software can expose the remote GPU through an API, SDK or virtual GPU driver, allowing applications to access acceleration without necessarily being rewritten for WiCi’s hardware.
That architecture could also change how GPU memory is viewed. WiCi uses NVMe storage as another tier in the memory hierarchy, allowing portions of large AI models to be streamed from storage rather than requiring the entire model to fit into GPU memory.
WiCi is targeting more than AI inference. Potential applications include gaming, image and video creation, voice assistants, autonomous agents, robotics and other workloads requiring substantially more compute than a client device can conveniently provide.
The company also sees potential in future AR and VR systems, where moving GPU hardware away from a wearable device could reduce weight, heat and power consumption.
The company says its RTX 5090-based system can run very large models, including the 2.8-trillion-parameter Kimi K3, at approximately five tokens per second. Those figures are vendor claims and will need to be evaluated through independent testing as the hardware becomes available.
WiCi was founded by researchers and consumer-hardware executives, with CEO and co-founder Zili Meng bringing a research background in electronic and computer engineering. The company says its research team has published at major computer-systems conferences and that its hardware team has experience shipping more than 30 million consumer devices.
The company plans to make a developer preview available in the fourth quarter of 2026, with WiCi One advertised at an introductory price starting at $1,999.




