I'm Zack Li (Zhiyuan Li), an entrepreneur, builder, and lifelong learner building scalable AI systems from frontier research to robust products. My interests lie in scalable software architecture design, Gen AI inference and Physical AI.
I'm currently a Senior Staff Machine Learning Engineer at Qualcomm, where our team is building GenieX, a generative AI inference runtime for Qualcomm platforms. I also contribute to llama.cpp for the Snapdragon Hexagon NPU, and to ai-hub-models, Qualcomm AI Hub's zoo of models optimized for on-device deployment.
I co-founded Nexa AI (acquired by Qualcomm) as Chief Technology Officer. On the engineering side, we build Nexa SDK (now GenieX), which reached #1 on GitHub Trending and #1 Product of the Day, earned 8k+ GitHub stars, and was featured by Qualcomm, AMD, NVIDIA, Intel, and IBM. On the research side, I co-authored the Octopus model series, which trended top-2 on Hugging Face and was featured by Google I/O and Google DeepMind. On the business side, I led enterprise relationships and customer delivery for HP, Lenovo, and İşbank, and technical showcases at CES 2024–2026, Snapdragon Summit, Microsoft Ignite, PyTorch Conference, and more.
Before Nexa, I built on-device AI and software systems at Google, Amazon Lab126, and Cadence. I hold an M.S. from Stanford University (2019) and a B.E. from Tongji University (2017).
GenAI inference runtime for NPU, GPU, CPU
Run any frontier model on the edge in one line of code — a single runtime across NPU, GPU, and CPU. Hit #1 on GitHub Trending and #1 Product of the Day.


On-Device Agentic AI Models
A pioneering agent model series using functional tokens for strong function calling across phones, PCs, wearables, cars, and robots. Reached Hugging Face's top three and spotlighted at Google I/O.


Brought our stack to Snapdragon for Android, the Hexagon NPU, Granite 4.0 at the edge, and IoT & robotics — plus seven mentions in "This Week in AI."
Teamed up to ship a local AI agent on RTX, showcased us as a featured startup at their CES 2026 booth, and ran a 1.3M-view feature on YouTube.
Named us a partner for CES 2026 and documented our image generation on its NPU and DeepSeek R1 speedups.
Spotlighted our work on its developer channel.
Included OmniAudio in the Gemmaverse and highlighted our work at Google I/O 2024.
Listed NexaML as a supported framework for Granite 4.0, next to vLLM, llama.cpp, and MLX.
Called out our day-0 Qwen3-VL support and Qwen3 on the Qualcomm NPU.
Teamed up with us for the LFM2.5 launch and LFM2.5 Thinking.
Thomas Wolf ranked our models among the most-downloaded, and CEO Clément Delangue made a public call to collaborate.
Amplified our work through Developer Relations.
Brought us on stage at the Ignite 2025 keynote as an official partner.