Behind the intelligence.
Original introductory notes on the systems that surround an AI model.

01 / ARCHITECTURETwo API servers, two different jobs.
Read note ⌄
A Kubernetes API server manages the desired state of a cluster. An inference API server receives application requests. Keeping these roles distinct makes the architecture easier to explain and operate.
A mobile prompt travels through the application path to a serving workload. Cluster configuration travels through the control plane. Both matter, but they solve different problems.
Primary reference: Kubernetes components ↗02 / PERFORMANCEA token is not a network packet.
Read note ⌄
A token is a unit used by the model and its text encoding. A network packet is a transport unit used to move data between machines. A streamed response can contain several tokens, and its underlying transport can split or combine data.
Chip links exchange computational data inside a system. A network fabric moves data between systems. A helpful diagram should show both without treating them as interchangeable.
03 / PROTOTYPESWhy a prototype should show its boundaries.
Read note ⌄
Simulated resources and mock providers make it possible to develop a workflow before live integrations are available. They help validate interaction, ranking, authorization and state handling.
Silverleaf’s food-comparison assistant currently uses mock providers, and Cloud Cost Intelligence uses simulated AWS/Azure resources. Making those boundaries visible helps visitors understand the current scope.

