Back to the profile

ENGINEERING EXPERIENCE

The project record.

Architecture, implementation, and technical leadership across data, ML, infrastructure, and industrial systems.

Platform Architect

US Digital Infrastructure Company

GPU compute platform

Architected and built GPU compute infrastructure supporting virtual machines and physical servers. Owned system design and hands-on implementation across resource provisioning, management, and remote access.

Compute management and automation

Built compute-management software using open-source virtualization, a web management console, and infrastructure inventory tooling. Developed automation across compute, hardware, and network operations.

AI application platform

Lead technical delivery of an AI application platform with engineering and product. Built isolated execution environments and service integrations while overseeing the agent core and data subsystem.

Access controls and observability

Established platform access policies and identity integrations. Implemented telemetry, monitoring, and auditing for platform operations.

Datacenter monitoring

Built monitoring and selected equipment controls for datacenter facility systems. Integrated physical infrastructure with platform software and external operational systems.

Platform security architecture

Designed and led implementation of security architecture for a multi-tenant platform. Covered identity, authorization, workload isolation, and data protection across services and infrastructure.

Hardware-assisted networking prototype

Built an initial prototype for distributed firewall and overlay-network functions using hardware acceleration. Connected platform software with programmable networking hardware.

LLM inference and routing

Integrated an open-source inference engine and designed and built request routing and load balancing for language-model workloads.

Sr. Principal Engineer, Platform Architect

US Agricultural Technology Company

Distributed ML training

Built distributed computer-vision training workflows across cloud and on-premises environments. Integrated different training systems, networking, and workload placement with shared orchestration and developer tooling.

Platform engineering leadership

Owned technical direction, development, and operations for a platform infrastructure and reliability team supporting data and ML engineers. Established on-call, service objectives, monitoring, and support practices.

Data-platform architecture and migration

Re-architected shared analytics software into an autoscaling, multi-tenant platform. Migrated large data collections while preserving compatibility and service continuity for existing workflows.

Self-service infrastructure

Built self-service cloud infrastructure so engineering teams could provision and adapt their own environments. Designed storage for analytical information and files.

Hardware-in-the-loop platform

Developed a research prototype into a hardware-in-the-loop testing platform. Automated equipment allocation, recovery, monitoring, and simulated inputs for application and system-software validation.

Identity and authorization

Designed and built shared identity and authorization services for people, applications, and devices. Connected engineering tools with centralized access controls and independent device-access management.

Infrastructure cost reporting

Made shared compute and data-infrastructure costs attributable to engineering teams. Delivered dashboards, budget alerts, and trend notifications to support engineering-group budgeting.

Infrastructure Tech Lead

US EV Manufacturer (through EPAM Systems)

Data and ML platform

Led architecture and engineering across a joint client and EPAM team. Built data-ingestion and model-workflow blueprints, implemented selected workflows, and integrated data-quality monitoring.

Infrastructure SDK and CLI

Created a shared infrastructure SDK and CLI for reusable provisioning and deployment components. Improved pipeline development with common abstractions, workflow tooling, and validation before deployment.

Workflow operations

Adapted workload execution and scheduling to the constraints of managed cloud services. Added deployment validation, build and release tooling, and operational monitoring.

Infrastructure Tech Lead

US Credit Bureau (through EPAM Systems)

Analytics platform leadership

Led engineering for internal and customer-facing analytics platforms from implementation through maintenance. Helped build the engineering team and owned technology decisions with product.

Platform architecture and reliability

Owned architecture and implementation of a cloud analytics platform, including redundant infrastructure, interactive development environments, automated provisioning, and service recovery.

Security design and assurance

Led platform security design and data-protection work. Supported internal certification, independent security audits, and compliance documentation.

Lead Systems Engineer

EPAM Systems

Healthcare ML platform

Led implementation of a Healthcare ML platform through production delivery. Coordinated engineers, an architect, and the customer on data-processing and model workflows.

Open-source ML platform

Built and released open-source ML training and deployment software. Connected notebook-based development with training infrastructure, model export, and deployment workflows.

Python engineering practice

Helped establish Python engineering within the data-engineering practice. Developed technical training for external participants, including curriculum, assessments, and mentoring; graduates joined the company.

Senior C++ Developer, Architect

Russian University Research Laboratory

Arctic monitoring application

Owned architecture and end-to-end development of an offline Arctic monitoring application. Integrated radar, vessel positions, and satellite imagery with data processing, databases, and visualization for remote field use.

Field deployment

Prepared software, infrastructure, and installation guidance for environments with limited connectivity and maintenance access. Tested deployment and failure scenarios and supported field installation, with attention to reliable operation and usability.

Industrial simulation

Built networking, communication, and deployment layers for industrial training simulators. Modernized a C++ environmental-simulation core, added capabilities, and automated testing.

Web Developer

Crocus City Hall

Ticketing and reservations

Delivered and supported a full-stack online ticketing and reservation platform for a concert venue. Owned the customer interface, backend, and booking workflows, starting the commercial project while in high school.

PROFILE

Education

Education and professional development

Earned a B.S. in Computer Science from ITMO University in 2016 and an M.S. with honors in 2018. Final projects covered industrial remote control and Arctic monitoring. Completed the Architecting with Google Cloud Platform Specialization in 2019 and earned the AWS Certified Machine Learning – Specialty certification in 2021.

PROFILE

Background & skills

Profile and contact

Kirill Makhonin is a principal software engineer and platform architect based in New York, NY, with 17+ years of experience. Contact: career@makhonin.biz. LinkedIn: https://www.linkedin.com/in/kirillmakhonin/. English: fluent. Russian: native.

Technical skills

Technical skills: Languages: Python, Go, TypeScript/React, C/C++, C#/.NET, Rust. Infrastructure: Kubernetes, Kubernetes webhooks & operators / Operator SDK, AWS, GCP, Terraform, Terraform providers, Terragrunt, Pulumi, Ansible, Cloud Hypervisor, NetBox, HA, Load Balancing, Linux, MacOS, HELM, NVidia GPU Operator, PXE/iPXE, Multus, IPMI, Redfish, Modbus. ML and data: Kubeflow, Slurm, EFA/RDMA, Airflow, Databricks, JupyterHub/Lab, Monte Carlo, Kafka, Redis, PostgreSQL, Graph Data Bases, vLLM, MLflow, Claude Code, MCP (development and consuming). Identity and AI: OIDC, mTLS/PKI, JWT/JWKS, OPA, Firecracker, MCP. Operations: Observability, Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, Loki, PagerDuty, Rootly, GitHub Actions, Jenkins, Agentic Development, Agentic Based Triaging and Operations. Networking: DHCP, DNS, Rest, ConnectRPC, gRPC, MetalLB, Traefik, Istio/Envoy. Security: PKI / HSM, JWT, OAuth/OIDC, Open Policy Agent, OpenBAO, TLS/mTLS.

Engineering focus

Hands-on software engineer and platform architect building distributed data, ML/AI, and infrastructure platforms across heterogeneous environments. Use agentic development workflows, primarily Claude Code, across architecture and codebase exploration, implementation, testing, and review, while retaining ownership of technical decisions and production quality. Build production systems with Python, Go, TypeScript/React, and a broad range of IaaC tooling.

17+ years of experience spanning industrial simulation and geospatial applications, enterprise data platforms, distributed ML training and data platforms, identity systems, GPU infrastructure, and AI application platforms.

Focus: distributed systems, private clouds, AI/ML and data workloads, and custom hardware integrations.

Contact Kirill