Resume

Vishal Reddy Kolan

Senior DevOps Engineer / Platform Engineer

9+ years of experience building cloud platforms, Kubernetes environments, infrastructure automation, CI/CD pipelines, and enterprise observability systems across AWS, Google Cloud Platform, and Azure.

Experience 9+ Years

Cloud, DevOps, release engineering, and platform operations.

Kubernetes 100+ Clusters

Operated multi-tenant fleets across GKE, EKS, AKS, and on-prem.

Clouds 3 Major Platforms

AWS, Google Cloud Platform, and Microsoft Azure.

Strengths CI/CD + IaC + Observability

Automation, reliability, governance, and production support.

Overview

Summary

Senior DevOps and Cloud Engineer with strong experience in software integration, configuration, packaging, automation, release engineering, and production deployments on Unix and Linux platforms. Experienced in supporting large-scale enterprise organizations with expertise in AWS, GCP, Azure, Kubernetes, Terraform, CI/CD automation, cloud networking, observability, and platform engineering.

Strengths

Professional Summary

  • Extensively worked on Jenkins, Screwdriver, and GitHub Actions for CI and end-to-end build and deployment automation.
  • Expertise in Terraform, CloudFormation, and ARM Templates to provision and manage infrastructure across AWS, GCP, and Azure.
  • Experience designing and managing Kubernetes clusters across AWS EKS, Google GKE, Azure AKS, and on-prem environments.
  • Strong hands-on experience in Docker, Kubernetes, Helm, and advanced Helm templating for multi-environment and multi-tenant deployments.
  • Implemented CI/CD frameworks integrating Jenkins, Maven, Docker, Helm, and Artifactory in Linux environments.
  • Worked deeply with AWS services such as EC2, VPC, S3, IAM, RDS, ELB, ALB, NLB, Route53, CloudWatch, SNS, SQS, and Auto Scaling.
  • Hands-on experience with GCP services including GKE, IAM, VPC, Cloud Monitoring, BigQuery, Private Service Connect, and Looker Studio.
  • Worked with Azure services including AKS, VNets, NSGs, Azure Data Factory, Databricks, Log Analytics, and Azure Monitor.
  • Implemented Istio service mesh for mTLS, traffic routing, retries, circuit breaking, and telemetry collection.
  • Strong Linux system administration and troubleshooting background.
  • Experience designing Kafka and event-driven architectures for scalable distributed systems.
  • Built observability platforms using Prometheus, PromQL, Grafana, ELK Stack, Datadog, and Splunk.
  • Developed automation scripts using Python and Bash for provisioning, reporting, monitoring, and operational efficiency.
  • Worked in Agile and Scrum delivery models with production support and incident management responsibilities.
  • Strong troubleshooting, analytical, communication, and problem-solving skills.

Toolkit

Technical Skills

Operating Systems

Linux (RHEL, CentOS, Ubuntu), Unix

Cloud Platforms

Amazon Web Services, Google Cloud Platform, Microsoft Azure

Networking Protocols

TCP/IP, FTP, HTTP, HTTPS, DHCP, DNS, NIS, NIS+, NFS

Containers & Orchestration

Docker, Kubernetes (EKS, GKE, AKS), Helm, Istio

Infrastructure as Code

Terraform, Terragrunt, CloudFormation, ARM Templates

Version Control

Git, GitHub, GitLab, SVN, Atlassian, Bitbucket

Scripting Languages

Bash, Python, Perl, Ruby, XML, JSON

Automation Tools

Ansible, Chef, Puppet

CI/CD Tools

Jenkins, Screwdriver, GitHub Actions, GitLab CI, Maven, Nexus, Artifactory

Databases

MySQL, PostgreSQL, BigQuery, Redshift

Monitoring Tools

Prometheus, Grafana, Datadog, ELK Stack, CloudWatch, Splunk

Cloud & Virtualization

AWS, Docker, Kubernetes, Azure

Career

Professional Experience

Yahoo

Senior DevOps / Platform Engineer

Nov 2022 - Present

Migrated Yahoo Mail's mailbox platform and supporting microservices from legacy on-premises Kubernetes clusters to cloud-native environments on Google Kubernetes Engine and AWS EKS, improving scalability, reliability, security, and operational efficiency.

  • Owned day-to-day operations of a large multi-tenant Kubernetes fleet with more than 100 clusters across multiple regions.
  • Designed, built, and operated Kubernetes platforms across GKE, AWS EKS, and legacy on-prem Kubernetes clusters.
  • Built AWS EKS clusters from scratch including VPC CIDR planning, public and private subnets, route tables, internet gateways, NAT gateways, security groups, network ACLs, ALB, NLB, Route53, and private DNS.
  • Supported on-prem Kubernetes clusters during phased migration to GKE and EKS including networking, certificates, DNS, ingress, and workload readiness.
  • Implemented and operated Istio service mesh across cloud and on-prem clusters enabling mTLS, traffic routing, canary deployments, retries, circuit breaking, and telemetry collection.
  • Designed shared VPC and tenant-project architectures with centralized networking and governance.
  • Provisioned and managed GCP IAM roles, bindings, and service accounts across shared and tenant projects using reusable Terraform modules and Git-based pull request workflows.
  • Implemented least-privilege access patterns for application onboarding including runtime identities, human access controls, and controlled break-glass access for production support.
  • Supported enterprise identity integrations and group-to-role mappings by aligning Google Cloud IAM, Kubernetes RBAC, and Athenz-based access standards.
  • Authored reusable Terraform modules for clusters, node pools, firewall rules, VPC peering, Private Service Connect, Private Service Access, private DNS zones, and monitoring resources.
  • Produced audit-ready onboarding evidence including access approvals, Terraform deployment history, and Cloud Audit Log references while maintaining operational runbooks and support documentation.
  • Built Helm charts and advanced Helm templating workflows supporting multi-tenant deployments, region overrides, namespace labels, and safe rollout strategies.
  • Implemented CI/CD pipelines using Screwdriver to manage infrastructure and platform releases with automated validation, promotion, and rollback.
  • Led compute and node pool migrations, handled quotas, reservations, autoscaler behavior, PDB coordination, and cost optimization initiatives.
  • Built centralized observability standards using Google Managed Prometheus, PromQL dashboards, recording rules, and alerting policies.
  • Created BigQuery datasets and SQL pipelines to analyze node images, upgrades, and workload metadata.
  • Built Looker Studio dashboards for capacity planning, upgrade readiness, reliability trends, and workload utilization.
  • Integrated Athenz identity, SIA certificates, RBAC, and IAM for secure user and service access.
  • Resolved IAM and access-related issues involving policy inheritance, service account permissions, cluster authentication, and cross-team dependencies with platform, network, and security teams.
  • Troubleshot complex production issues including API latency, NAT saturation, inter-cluster networking, and certificate-based health checks.
  • Integrated Splunk for centralized log ingestion, audit visibility, and production troubleshooting across Kubernetes and infrastructure layers.
  • Worked with Google Cloud Support and product teams for advanced troubleshooting and feature enhancement requests related to GKE and Cloud Monitoring.
  • Produced runbooks, automation scripts, and internal tooling to standardize audits and operations.
  • Acted as the primary on-call contractor handling incidents, RCA, scaling tests, and capacity planning reviews.

Environment: AWS, GCP, Kubernetes (GKE/EKS), Istio, Terraform, Helm, Prometheus, Google Cloud Monitoring, Splunk, BigQuery, Looker Studio, Python, Bash

Dell

DevOps Engineer

Aug 2021 - Nov 2022

Built an Azure Databricks-based data platform to migrate large-scale telemetry workloads from AWS S3 to Azure Data Lake, with secure networking, infrastructure automation, CI/CD, and enterprise analytics support.

  • Delivered cloud automation and platform engineering solutions supporting enterprise telemetry processing and analytics workloads.
  • Designed end-to-end Azure infrastructure including resource groups, VNets, subnets, NSGs, route tables, private endpoints, service endpoints, and secure connectivity to AWS S3 sources.
  • Automated Azure infrastructure provisioning using Terraform and ARM Templates across development, staging, and production environments.
  • Designed AKS cluster architecture with node pools, autoscaling policies, ingress configuration, and secure network integration with Azure services.
  • Deployed and operated containerized microservices on AKS using Docker, Helm, and CI/CD pipelines with blue-green and rolling deployment strategies.
  • Built and maintained Azure Data Factory pipelines to ingest telemetry data from AWS S3 into Azure Data Lake Storage Gen2.
  • Integrated Azure Databricks Spark jobs with ADF for large-scale batch and near real-time processing workloads.
  • Implemented CI/CD pipelines for infrastructure and application deployments using Jenkins, Maven, Docker, GitLab, Bitbucket, and Artifactory.
  • Standardized Docker image builds using Docker-Maven plugins and centralized artifact management.
  • Managed PostgreSQL flexible servers including provisioning, high availability configuration, backup strategies, and performance tuning.
  • Implemented centralized logging, metrics, and alerting using Azure Log Analytics and Azure Monitor.
  • Designed secure identity and access management controls using Azure AD roles and RBAC policies.
  • Supported production deployments, coordinated releases, and performed root cause analysis for performance and connectivity issues.
  • Collaborated with data engineering and analytics teams to optimize Spark job performance, cluster sizing, and cost efficiency.

Environment: Azure, AKS, Terraform, ARM Templates, Azure Data Factory, Azure Databricks, Docker, Jenkins, GitLab, Maven, PostgreSQL, Azure Monitor, Log Analytics

Qubole

DevOps Engineer

Jan 2020 - Aug 2021

Standardized enterprise Kubernetes platforms across AWS EKS and Google GKE, migrated legacy CloudFormation stacks to Terraform and Terragrunt, and improved delivery, governance, and observability for financial platforms.

  • Led Kubernetes platform standardization initiatives across AWS EKS and Google GKE clusters.
  • Designed and provisioned Kubernetes clusters including node groups, autoscaling policies, IAM roles, network configurations, and security boundaries.
  • Migrated AWS CloudFormation stacks to Terraform and Terragrunt, improving infrastructure reusability, governance, and multi-environment consistency.
  • Implemented remote state management, module versioning, and environment-based configuration patterns for Terraform deployments.
  • Built reusable and parameterized Helm charts to support multiple microservices across environments.
  • Managed Kubernetes manifests including deployments, services, ingress resources, config maps, secrets, and resource quotas.
  • Implemented CI/CD pipelines using Jenkins, GitHub, GitLab, and Spinnaker to automate build, test, containerization, and deployment workflows.
  • Automated Jenkins administration including plugin lifecycle management, pipeline standardization, access control, and security hardening.
  • Implemented container security best practices including image scanning, RBAC enforcement, and namespace-level isolation.
  • Deployed and managed Elastic Cloud Enterprise using Terraform, Ansible, and shell scripting for scalable logging infrastructure.
  • Implemented Istio service mesh for secure service-to-service communication, traffic routing, retries, and observability.
  • Configured NGINX ingress controllers and TLS termination for secure external access to services.
  • Built Kibana dashboards, watcher alerts, and Logstash pipelines using Grok patterns to parse application and infrastructure logs.
  • Integrated Kubernetes logs, application logs, and VPC Flow Logs into centralized ELK for audit, troubleshooting, and monitoring.
  • Implemented monitoring using Prometheus and Grafana for cluster health, node performance, and workload visibility.
  • Performed production troubleshooting including networking issues, pod scheduling failures, resource constraints, and scaling inefficiencies.
  • Participated in incident response, RCA, and post-incident reviews to improve platform reliability.

Environment: AWS, GCP, EKS, GKE, Kubernetes, Terraform, Terragrunt, Helm, Jenkins, GitHub, GitLab, Spinnaker, Istio, NGINX, ELK Stack, Ansible, Docker, Linux, Python

UAE Exchange

DevOps Engineer

Oct 2017 - Dec 2019

Supported enterprise financial and payment platforms with CI/CD automation, Kubernetes modernization, centralized logging, GCP monitoring, and 24x7 production support.

  • Supported enterprise-grade financial transaction systems requiring high availability, security, and regulatory compliance.
  • Built and maintained Jenkins-based CI/CD pipelines for Java, Spring Boot, and Node.js applications including automated build, test, artifact publishing, and deployment stages.
  • Integrated source control systems such as Git and Bitbucket with Jenkins and implemented branching strategies to support Agile development workflows.
  • Containerized monolithic applications using Docker and migrated workloads to Kubernetes to improve scalability and deployment consistency.
  • Provisioned and operated Kubernetes clusters including node configuration, resource quotas, ingress setup, TLS configuration, and horizontal pod autoscaling.
  • Implemented secure networking policies and RBAC within Kubernetes environments.
  • Configured load balancers and ingress controllers to securely expose financial APIs and internal services.
  • Implemented GCP Stackdriver dashboards, custom log-based metrics, and alerting policies to monitor application health and system performance.
  • Centralized logging across 5,000+ servers using Splunk forwarders, indexers, and search heads for traceability and audit compliance.
  • Designed and maintained log retention, parsing, and indexing strategies for compliance and operational analysis.
  • Managed GCP Pub/Sub topics, logging sinks, and access controls for secure log routing and data handling.
  • Automated routine operational tasks using Bash scripting to reduce manual intervention and improve reliability.
  • Participated in 24x7 on-call rotations, handled production incidents, performed RCA, and implemented preventive measures.
  • Collaborated with development and QA teams to adopt CI/CD best practices, container standards, and Kubernetes deployment patterns.

Environment: GCP, Kubernetes, Docker, Jenkins, Splunk, Stackdriver, Git, Bash

A3IT Solutions

Build & Release Engineer

Oct 2016 - Sep 2017

Provided build and release engineering support, Linux infrastructure administration, deployment automation, and operational documentation for enterprise client environments.

  • Supported build and release processes for enterprise web applications across development, QA, and production environments.
  • Managed Linux infrastructure including OS installation, patching, kernel updates, security hardening, and performance tuning.
  • Automated build and deployment workflows using Bash and shell scripts to reduce manual effort and improve reliability.
  • Scheduled automated jobs using cron for deployments, backups, and routine maintenance.
  • Installed, configured, and maintained Apache HTTP Server and Tomcat application servers for Java-based applications.
  • Managed user accounts, file permissions, SSH access, and basic security configurations across Linux servers.
  • Supported enterprise networking services including DNS, DHCP, NFS, LDAP, and SSH connectivity.
  • Assisted in troubleshooting application startup failures, server performance issues, and configuration errors.
  • Maintained artifact repositories and supported build tools for packaging and releasing application binaries.
  • Documented deployment procedures, rollback steps, troubleshooting guides, and environment setup instructions.
  • Coordinated with development and QA teams to ensure successful release cycles and timely deployments.

Environment: Git, Linux, Apache, Tomcat, Bash, Cron, SSH, DNS, DHCP, NFS, LDAP, VMware, SMTP, FTP, Python, Ruby, JBoss, Nginx