Writing & appearances
Where I've written, spoken, taught, and rambled about AI.
Articles
-
NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents
-
Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard
-
Start Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in Minutes
-
NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents
-
Building NVIDIA Nemotron 3 Agents for Reasoning, Multimodal RAG, Voice, and Safety
-
Introducing Nemotron 3 Super: An Open Hybrid Mamba-Transformer MoE for Agentic Reasoning
-
How to Train an AI Agent for Command-Line Tasks with Synthetic Data and Reinforcement Learning
-
How to Build a Voice Agent with RAG and Safety Guardrails
-
Inside NVIDIA Nemotron 3: Techniques, Tools, and Data That Make It Efficient and Accurate
-
Develop Specialized AI Agents with New NVIDIA Nemotron Vision, RAG, and Guardrail Models
-
Build More Accurate and Efficient AI Agents with the New NVIDIA Llama Nemotron Super v1.5
-
Build an AI Agent with Expert Reasoning Capabilities Using the DeepSeek-R1 NIM
-
Mastering LLM Techniques: Evaluation
-
Deploying Fine-Tuned AI Models with NVIDIA NIM
Book
Research
-
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
550B-parameter MoE hybrid Mamba-Transformer with 1M-token context, built for high-throughput agentic reasoning.
-
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
120B (12B active) LatentMoE hybrid Mamba-Attention model — first in the family pre-trained in NVFP4, with MTP layers for native speculative decoding.
-
NVIDIA Nemotron 3: Efficient and Open Intelligence
Technical report for the Nemotron 3 family — MoE hybrid Mamba-Transformer models (Nano, Super, Ultra) with NVFP4 training, LatentMoE, and 1M-token context.
-
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
-
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
-
Llama-Nemotron: Efficient Reasoning Models
Talks
Podcasts
Videos & Courses
-
Nemotron Labs Livestream
I host conversations and Q&As with the people building Nemotron on the NVIDIA AI channel.
-
The AI Engineering Bootcamp
The bootcamp I co-created and taught: prompt engineering, RAG, fine-tuning, agents, and evals. The curriculum became our book (Wiley, 2026).
-
Chris Alexiuk on YouTube
Videos on LLMs, fine-tuning (LoRA and friends), and ML techniques.
-
AI Makerspace on YouTube
A back catalog of live sessions I co-hosted on LLMs, agents, evals, and AI engineering tools.