HUM GEN AI NA Naoki Goto Understanding Attention From Embeddings to KV Cache: How Transformers Actually Work
TCH AI NA Naoki Goto From 2s to 600ms: PyTorch vs ONNX Runtime Benchmarking DistilBERT Inference on EKS
TCH MDA AI NA Naoki Goto Deploying & Observing Qwen2 on AWS EKS Serving Qwen2–1.5B on AWS EKS with GPU: from infrastructure setup to VRAM observation
TCH AI NA Naoki Goto Deploying DistilBERT on AWS EKS From Minikube to AWS EKS: Infrastructure as Code, Load Testing, and What CPU Throttling Does to Inference Latency
TCH NA Naoki Goto Raft: Part 2 (Log Replication) A developer’s perspective on Raft log replication: AppendEntries RPC, commitIndex vs lastApplied, and the Figure 8 problem explained.
TCH NA Naoki Goto Raft in Kubernetes: etcd Deploying a Go API on Kubernetes and peeking inside etcd to see where Raft lives in production.
TCH AI NA Naoki Goto Deploying DistilBERT on Kubernetes What I Learned About Resource Limits and Inference Latency
TCH NA Naoki Goto Implementing CI/CD Pipeline with GitHub Actions Streamlining the deployment and eliminating human errors
TCH NA Naoki Goto Defense in Depth: Building a Resilient API with Go and Kubernetes What I learned about reliability and security by building a Go API and deploying it on Kubernetes.
TCH AI NA Naoki Goto Containerizing DistilBERT A beginner’s guide to deploying an open source NLP model with Docker