| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this userβs behavior. Learn more about reporting abuse.
Report abuseInfrastructure engineer with 12+ years of experience, including time at Amazon Web Services, now focused on the operational backbone behind large-scale AI/ML systems β Kubernetes, GPU orchestration, and the reliability tooling that keeps model-serving infrastructure observable, debuggable, and recoverable at scale.
"End-to-end Terraform stack for GPU workloads on EKS β VPC, cluster, and a spot/on-demand GPU node pool."
HCL
Kubernetes operator for GPU node health β cordons/drains nodes on Xid/ECC errors or thermal throttling, zero external dependencies.
Go
Helm chart for deploying vLLM on GPU Kubernetes β startup probes tuned for slow model loads, correct GPU scheduling, sized /dev/shm, and GPU-utilization/queue-depth autoscaling.
Go Template
| Back | FazBrowse Home | New Git URL |