Collin
Hargreaves

UC Berkeley

interests →machine learning, inference engineering, distributed systems, systems programming, cloud infrastructure, startups
hobbies →djing, piano, cars, reading, sports, traveling

the best soda is cherry coke zero

Hey, my name is Collin, I'm currently a Junior at UC Berkeley. Before that, I was in the Air Force working in Aerospace Propulsion.

I have previous experience working on several AI/ML SageMaker teams at AWS. I worked on the integration and launch of MLflow into Sagemaker, an MLOps tool. Along with that I built Inference load balancing for HyperPod's newly developed inference operator. Where I got to learn about many things, routing optimization, serving engines (sglang, vllm), latency and throughput anaylsis, kuberenetes, and cluster orchestration. My work was validated and pushed to production-scale traffic after launch.

I am now currently working on the HyperPod Dataplane team where I am working on a project involving the decoupling of machine learning images from agent infrastructure to simplify cluster redeployments.

Apart from work, I also love building. Right now I'm working on agentic operations for engine test cell and satellite operations. I'm also learning how to build an LLM from scratch, Rust, plus how to make house music and DJ, some of my favorite artists are KETTAMA, MallGrab, John Summit and i_o.

Sponsorship Director at Cal Hacks. Reach out at collin@hackberkeley.com.

projects

Interested in building at the infrastructure level

Yuliademo / mvp

A demo for AI agents in aerospace operations. I simulated a rocket launch in Kerbal Space Program and piped its live flight data through kRPC into Foundry, where the vehicle is modeled as ontology objects. An agent named PROP watches the engines like a propulsion flight controller would, catches engine-out failures, and proposes a recovery command. The flight dynamics math for thrust-to-weight, delta-v, and orbit reachability is computed deterministically, so the agent decides what's wrong while the physics stays exact. A human approves every command before it reaches the vehicle. Point the same pipeline at a real telemetry feed and nothing downstream changes.

PythonFoundryAIP LogickRPCOntology
github →
mainstacks

Like a producer reusing loops they have made before, mainstacks lets you extract patterns from past projects and drop them into new ones. Your agents get context on how you build things, so they stop guessing and start building the way you would

GoDevtoolHomebrew
github →
Keel

Wrap your LLM client with one line of Python and every call, token, cost, and source location shows up in your dashboard in real time. Built to make agentic AI costs visible per agent, per function, and per line of code before they surprise you.

PythonFast-APIMongoDBMCPSDKNext.jsTailwind
github →
TreeAgent - Accepted to ICML 2026paper ↗

Co-authored a paper accepted to ICML 2026 AI for Science Workshop: developed a multi-agent system (MAS) that orchestrates expert decision trees with Vision-Language Models (VLMs) for automated bias labeling in forestry remote sensing, outperforming supervised ML baselines while preserving interpretability

PythonLangChainLangGraphVLMPDALraster
Recepta

Automates medical referral intake, OCR to structured extraction to EMR auto-fill. Taught me more about customer discovery and what people actually pay for in the context of building a startup.

Next.jsFastAPITesseractBrowser Agents
github →
nvim-config

My Neovim configuration. Clean, fast setup focused on development with LSP support, fuzzy finding, and Treesitter. Uses Packer.

LuaLSPTelescopeTreesitterMason
github →
NoBias

One of my first projects, built with some friends at the first AI Hackathon from UC Berkeley. Analyzes news article bias with GPT-4 scoring and Hume AI sentiment, then rewrites the piece from three perspectives.

ReactPythonFastAPIGPT-4Hume AI

in progress and learning

tinykvin progress

Working through Sebastian Raschka's Build a Large Language Model From Scratch, then building tinykv on top of it. A minimal transformer inference engine in C++ focused on the KV cache manager, the scarce resource that determines throughput in distributed LLM serving.

C++

readings

I really like to read about things that show new ways to think about things.

4.6, amazing worldbuilding, and ideas of politics, religion, control, etc

3.9, great, but definitely not as good as the first

4.0, Again, but good

4.9, probably my favorite book

3.2, fun story, nothing too crazy

3.5, remember it being decent

4.1/5 Loved this book, there were some moments where the science dragged on a bit, but overall amazing

1st book 4.3/5, 2nd 3.9, 3rd 3.3, great series though

My childhood

WaldenHenry David Thoreau
Civil DisobedienceHenry David Thoreau
Self-RelianceRalph Waldo Emerson