About Us

CLOUDSUFI is a Silicon Valley-based specialist Data Engineering & Cloud Technologies player with top-tier clients, favorable revenue mix, strong financial performance, and robust management. We help organizations with Data Discovery, Insights, and Monetization, offering our engineers the opportunity to work on new platforms and technologies-including Cloud Hyper Scalers and AI/ML/NLP-that put them ahead of peers in the IT Services industry. Started in 2019, CLOUDSUFI is a family of 250+ members working towards a common goal of making enterprise data dance.

This role sits within CLOUDSUFI’s live engagement where AI is already embedded operationally across reliability and platform engineering-not a future aspiration. The team runs an AI SRE co-pilot that reasons across observability, cloud, and ticketing platforms; a set of purpose-built AI sub-agents for incident triage, monitor hygiene, and usage/cost attribution; and a formal AI governance program that assesses CLOUDSUFI’s own internal AI agents for risk before they touch production. You will be embedded with the client’s SRE function supporting a regulated consumer-lending platform on AWS-hundreds of microservices, event-driven pipelines, and customer-facing loan origination and servicing journeys where reliability is directly a customer-trust and compliance concern. As a Junior SRE Engineer, you will work under the guidance of senior engineers and architects to support, operate, and learn from this AI-driven reliability program.

Responsibilities - AI-Driven Initiatives


1. AI-Augmented Incident Response & Root Cause Analysis

  • Support incident response as a shadow or secondary responder, using AI-driven detection, correlation, and root-cause analysis tooling under the guidance of senior engineers to speed up triage.
  • Help assemble incident timelines from metrics, logs, traces, and deploy history, learning to separate the alert that fired from the change that actually caused it.
  • Help operate and monitor the existing AI SRE sub-agents (covering areas such as incident summarization, monitor-gap detection, and usage attribution), flagging anomalies or failures for senior review.
  • Assist in documenting AI-assisted incident findings and postmortems, including tracking corrective and preventive actions through to closure.

2. AI-Driven Observability & Alert Quality

  • Build and maintain monitors, dashboards, and SLO definitions in Datadog as code via Terraform, rather than clicking through the UI.
  • Assist in reducing alert noise using AI/ML-assisted detection (anomaly, outlier, and forecast monitors, composite conditions, and dynamic thresholds) so that pages stay actionable and rare.
  • Run recurring monitor-hygiene and coverage-gap reviews - no-data monitors, missing no-data notification, orphaned monitors on decommissioned services, and monitors with no owning team tag or valid notification target.
  • Support SLI/SLO and error-budget reviews for critical customer journeys, escalating services trending toward budget exhaustion to senior team members.

3. AI-Enabled Cloud Reliability, Capacity & Cost Engineering

  • Support day-to-day reliability of AWS workloads - ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS and Step Functions - under senior guidance.
  • Use predictive analytics and intelligent monitoring dashboards to spot capacity, saturation, and cost anomalies across host, container, log, APM, and serverless usage, escalating notable trends with evidence.
  • Assist in observability and cloud cost governance, attributing spend and telemetry volume to owning teams and services and helping identify low-value, high-cost signals.

4. AI Governance, Risk & Reliability Reviews (Learning Track)

  • Participate in architecture and design reviews, learning how reliability risk and AI-agent risk (e.g., prompt injection, authorization gaps, secrets exposure in agent and tool configuration, blast radius of autonomous actions) is assessed and documented.
  • Learn how production-readiness, change management, and audit expectations apply in a regulated environment (PCI-DSS, SOC 2, SOX, GLBA), and why observability evidence matters to auditors.
  • Assist in tracking action items from the AI-risk and reliability remediation roadmap and help prepare status updates for stakeholders.

5. AI-Enabled DevOps, CI/CD & Automation

  • Support the implementation of reliability and safety controls within CI/CD pipelines, learning automation-first and AI-augmented approaches and shift-left practices from senior engineers.
  • Help integrate AI-assisted checks (e.g., change-risk summaries, infrastructure drift detection, dependency and configuration checks) into the delivery lifecycle.
  • Write and maintain automation in Python and Bash against platform APIs (Datadog, AWS, GitHub, PagerDuty, Jira) to replace recurring manual work, and contribute Terraform modules and pull requests under review.
  • Assist with runbook creation and the progressive automation of runbook steps toward self-healing.

6. AI-Enabled Program & Workflow Support

  • Support multiple concurrent reliability initiatives using AI-enabled workflow and tracking tools shared across SRE and Security teams.
  • Help promote a reliability-first culture by learning and sharing data-driven, AI-backed operational practices with the wider team.

SRE & Cloud Capabilities

  • Interest in observability and monitoring fundamentals - metrics, logs, traces, and APM - and a willingness to learn predictive analytics and anomaly detection.
  • Exposure to at least one observability platform (Datadog preferred; Grafana/Prometheus, New Relic, CloudWatch, or ELK/OpenSearch also relevant).
  • Foundational understanding of SLI, SLO, and error-budget concepts, and of why alert fatigue is a reliability problem rather than an annoyance.
  • Hands-on exposure to AWS (or an equivalent hyperscaler) across compute, networking, storage, and managed database services, with interest in automated reliability tooling.
  • Foundational knowledge of container and orchestration concepts (Docker, ECS or Kubernetes) and of serverless execution models.
  • Beginner-to-intermediate familiarity with Infrastructure as Code (Terraform preferred; Ansible or CloudFormation acceptable) and an eagerness to build deeper IaC expertise.
  • Basic Linux troubleshooting and networking fundamentals, including DNS, TLS, load balancing, timeouts and retries.
  • Some hands-on exposure to automation or scripting (Python, Bash, or similar), including consuming REST APIs and parsing JSON.
  • Comfort with Git and pull-request-based workflows, and exposure to CI/CD tooling (GitHub Actions, Jenkins, GitLab CI, ArgoCD or similar).
  • Familiarity with incident management and on-call concepts, including severity models, escalation policies, and paging tools such as PagerDuty or Opsgenie.
  • Basic exposure to cloud posture and continuous monitoring concepts, and to change and release management discipline.
  • Practical curiosity about LLM-based assistants and agents applied to operations, and about how to verify whether their output is actually correct.

About You

  • 1–3 years of experience in SRE, DevOps, cloud infrastructure, platform, or production-support engineering, with an eagerness to build deeper expertise in infrastructure as code.
  • Some exposure to cloud-native logging and monitoring tools, and interest in AI-driven alerting and noise reduction.
  • You debug from evidence: when something breaks, your instinct is to look at the data before offering a theory, and you are comfortable saying “I don’t know yet, here is what I am checking.”
  • You would rather automate a task the second time you do it than the tenth.
  • Coursework or hands-on exposure to monitoring and observability platforms.
  • Exposure to CI/CD pipelines and interest in integrating reliability controls and shift-left practices.
  • Basic understanding of capacity, performance, and reliability cost concepts, including risk-based prioritization of remediation work.
  • Awareness of common compliance frameworks (PCI-DSS, SOC 2, SOX, HIPAA/GLBA) and comfort working with production-change discipline.
  • Willingness to learn incident management processes, including AI-assisted triaging, and the judgment to escalate early rather than sit quietly on an uncertain production signal.
  • Strong analytical and problem-solving skills, with curiosity to assess simple architectures for failure modes under senior guidance.
  • Clear written communication - you can explain an incident, a metric, or a trade-off to someone who was not in the room.
  • Interest in learning how reliability standards, operational policies, and governance frameworks are developed and maintained, including for AI-agent usage.

Core Competencies

Working knowledge of, or strong willingness to learn, several of the following domains:
  • Site Reliability Engineering fundamentals (SLI/SLO, error budgets, toil reduction)
  • Observability and Monitoring (metrics, logs, distributed tracing, APM, RUM/Synthetics)
  • AIOps basics - anomaly detection, event correlation, predictive analytics, alert noise reduction
  • Incident Response and Postmortem/RCA practice
  • Cloud Infrastructure (AWS core services, multi-account and IAM basics)
  • Containers and Orchestration (Docker, ECS, Kubernetes)
  • Infrastructure as Code and Configuration Management (Terraform, Ansible)
  • CI/CD, Release Engineering and Progressive Delivery
  • Automation and Scripting (Python, Bash, platform APIs)
  • Capacity, Performance and Reliability Cost Engineering (FinOps fundamentals)
  • Resilience Patterns (retries, timeouts, circuit breakers, graceful degradation)
  • Chaos/Failure-Injection and Disaster Recovery concepts
  • Runbook Automation, Automated Remediation and Self-Healing (exposure)
  • DevSecOps and secure-by-default operations fundamentals
  • AI Agent Risk & Governance basics (authorization models, prompt-injection risk, secrets exposure in AI tool configs, human-in-the-loop controls)

Preferred Certifications

  • AWS Certified Cloud Practitioner / Associate-level AWS certification (preferred, not required)
  • HashiCorp Certified: Terraform Associate (preferred, not required)
  • Datadog Fundamentals or an equivalent observability-platform certification (preferred, not required)
  • Certified Kubernetes Administrator (CKA) or KCNA (preferred, not required)
  • Exposure to AI/ML applications in operations and reliability is a plus (preferred, not required)

Required Skills

SRE Observability AI-reliability AI-Observability AI