Software Engineer · Private 5G & Distributed Systems
Building backend infra where failures are expensive —
Nokia · Starlink telemetry pipelines, billion-record ETL at Amazon,
and cloud-native systems that hold up under real production pressure.
I'm a Software Engineer currently at Armada AI in Bellevue, WA, working on Private 5G infrastructure. My work spans end-to-end data ingestion pipelines from Nokia and Starlink, distributed webhook delivery systems, SIM lifecycle management APIs, and full observability stacks.
Before that I spent 2.5 years at Amazon shipping high-throughput Java APIs, ETL pipelines at billion-record scale, and cloud-native microservices on AWS.
I care about systems that are observable, reliable, and built to last. I own production incidents, write the docs, and stick around after the merge.
Currently pursuing an MBA at WestCliff University, deepening my understanding of product strategy, business operations, and the intersection of technology and business decision-making.
Building production backend systems across startups and big tech
CS, Data Sciences, and MBA spanning India and the US
JWT, HMAC, Azure Key Vault, OAuth2 and CVE patching in production
Solutions Architect Associate with deep AWS and Azure production experience
End-to-end ingestion of Nokia NDAC data — RANs, SIMs, fault alarms, and 16 performance management counters — into Armada's unified P5G analytics platform. Built, broke, and hardened across real Nokia gNB hardware (gNB498730).
Nokia PM formulas come as raw strings like ETHIF_DL_BYTES / interval — built a recursive Compute() evaluator with row-level aggregation to prevent Cortex duplicate-series rejection.
IXR and HPE Edge alarms had no asset match and were silently dropped. Added NHG-based fallback — result: 0 skipped alarms.
Nokia SIM API paginates at 100 items. 94 of 102 SIMs were silently missing. Added automatic pagination loop — all SIMs ingest in one cycle.
Built the entire notification-svc from scratch — a production webhook delivery platform with HMAC signing, exponential backoff retry, async goroutine dispatch, ScyllaDB audit trail, and Mailgun failure alerting.
Slow customer endpoints couldn't block the pipeline. Moved all dispatch into goroutines with exponential backoff retry (3 attempts, 2-hour intervals).
Silent failures after max retries went unnoticed. Now marks webhook as ERROR in user-pref-svc and sends Mailgun alert with failure metrics to technical admins.
Prometheus panicked on restart from duplicate metric registration. Wrapped constructors in sync.Once — fixed CrashLoopBackOff across all restarts.
Full migration of SIM provisioning, deprovisioning, and editing APIs from v1 to v2 — cleaner REST shape, proper NDAC operation mapping, and handling of Nokia-specific edge cases that previously caused silent data loss.
Nokia NDAC returns PARTIAL_SUCCESS for valid operations. v1 treated it as error and rolled back. Now handled as success with a log entry.
After deprovision, NetworkName stayed in DB causing stale re-provisioning. v2 now clears it and updates UE_STATUS in the same transaction.
v1 full overwrites cleared unrelated fields. v2 uses targeted PATCH with partial DB updates — also had to add PATCH to Istio virtual service config.
Replaced ad hoc logging with a structured, unified observability layer across hyperwave-ingestion-svc, hyperwave-svc, and bff-layer — Prometheus metrics collection, structured slog standardization, and 5 Grafana dashboards covering every operational angle.
Inconsistent log.Printf across 3 services made incident debugging painful. Replaced with unified slog JSON logger with standard key-value fields.
5 Grafana dashboards: ingestion success rates, P99 latency, Go runtime health, Kubernetes pod metrics, and 16 Nokia PM instruments via a dedicated adapter.
Google Maps API–backed pipeline that converts raw Nokia lat/lng coordinates into human-readable addresses for 5G sites and RAN assets. Fixed Vault key loading so the Google API token is correctly injected via Azure Managed Identity.
Comprehensive audit trail across webhook lifecycle events (create, rename, delete, URL update) and SIM operations (provision, deprovision, edit) in user-audit-svc. Pre-deletion name capture ensures accurate records even after entity removal.
Armada AI
Bellevue, WA
Amazon Services LLC
Bellevue, WA
Uclid IT India Private Limited
India
AbhiBus Services
India
WestCliff University
United States
University at Buffalo, SUNY
Buffalo, NY
Vignan's Foundation for Science Technology and Research
India
Open to senior backend and distributed systems roles — especially infra, cloud-native, or telecom/5G. Reach out directly or book time below.
30 min · Google Meet
Krishna Charan
AI avatar · ask me anything