Key Responsibilities:
AI Platform / Kubernetes Application Engineer
- 24 years of professional experience in Kubernetes-based application development, platform engineering, DevOps, or cloud-native software engineering.
- Strong hands-on experience developing, containerizing, deploying, and operating applications on Kubernetes or OpenShift.
- Proficiency in Python and experience building production-grade APIs and microservices using frameworks such as FastAPI, Flask, or Django.
- Strong experience with Docker or Podman, including container image creation, multi-stage builds, image optimization, registry management, and container security practices.
- Experience creating and maintaining Kubernetes resources such as Deployments, StatefulSets, Services, Ingress, ConfigMaps, Secrets, Jobs, CronJobs, Persistent Volumes, and Network Policies.
- Hands-on experience with Helm, Kustomize, Kubernetes Operators, or similar deployment and configuration-management tools.
- Strong understanding of Kubernetes concepts, including pod lifecycle, scheduling, probes, resource requests and limits, autoscaling, storage, networking, RBAC, and service discovery.
- Experience troubleshooting application, container, networking, storage, and resource-related issues in Kubernetes environments.
- Practical experience with CI/CD tools such as GitLab CI, Jenkins, Argo CD, Tekton, or GitHub Actions.
- Experience implementing GitOps-based deployment and application lifecycle-management practices.
- Exposure to observability tools such as Prometheus, Grafana, Loki, Elasticsearch, OpenSearch, or distributed tracing platforms.
- Strong understanding of Linux, shell scripting, REST APIs, YAML, Git, networking fundamentals, and secure application configuration.
- Experience working with databases and data platforms such as PostgreSQL, MongoDB, Elasticsearch, ClickHouse, Redis, or vector databases.
- Experience deploying stateful and stateless applications in enterprise or air-gapped environments is highly desirable.
- Familiarity with secrets management, image vulnerability scanning, role-based access control, TLS certificates, and enterprise security controls.
- Experience with GPU-enabled Kubernetes workloads, NVIDIA GPU Operator, model-serving platforms, or GPU resource management is an advantage.
Secondary AI/ML Skills
- Basic to intermediate understanding of machine learning, deep learning, and generative AI concepts.
- Exposure to Python ML libraries such as Pandas, NumPy, scikit-learn, XGBoost, PyTorch, or TensorFlow.
- Familiarity with LLMs, transformers, Hugging Face, Ollama, RAG frameworks, MCP, LangChain, Haystack, CrewAI, or custom agent orchestration is preferred but not mandatory.
- Exposure to model serving, inference APIs, model quantization, MLflow, Weights & Biases, or other model and experiment-tracking tools is an added advantage.
- Familiarity with multimodal models and AI data pipelines is beneficial.