Data First Jobs

JPC TECHNO INC

Machine Learning Engineer II- Premium

Contract · In Office · Phoenix, Arizona (USA)

Posted Aug 4, 2026

  • We are seeking a highly skilled Senior Observability Operations Engineer to manage and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kuberetes, Linux, and cloud-native observability solutions.
  • Experience leveraging AI/ML and Generative Al to improve observability, automate operations, and accelerate incident resolution is highly desirable.
  • The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications.
  • Key Responsibilities
  • * Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
  • * Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
  • * Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
  • * Configure Dynatrace OneAgent, ActiveGate, Synthefic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis Al, and Application Performance Monitoring (APM).
  • * Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
  • * Develop dashboards; alerts, reports, and executive operational metrics.
  • * Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
  • * Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
  • * Perform root cause analysis for production incidents using observability platforms.
  • * Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
  • * Automate operational tasks using Python, Shell scripting, REST APls, Terraform, or Ansible.
  • * Participate in incident, problem, change, and release management processes.
  • * Drive platform upgrades, patching, security compliance, and operational governance.
  • * Improve platform reliability through automation, self-healing, and Al-assisted operations.
  • Required Technical Skills
  • Observability Platforms
  • * Dynatrace Administration
  • * Splunk Enterprise Administration
  • * OpenSearch Administration
  • * Elasticsearch Administration
  • * Grafana
  • * Prometheus
  • * Kibana
  • * Jaeger
  • * Open Telemetry
  • * Kafka (preferred)
  • Infrastructure
  • * Linux Administration
  • * Kubernetes
  • * Docker
  • * OpenShift or Rancher
  • * Networking (TCP/IP, DNS, Load Balancers, Firewalls)
  • * System Administration
  • Cloud & DevOps
  • * AWS, Azure, or GCP
  • * CI/CD pipelines
  • * Git
  • * Terraform
  • * Ansible
  • * REST APIs
  • Scripting
  • * Python
  • * Bash/Shell
  • * PowerShell (preferred)
  • AI & Automation Skills (Preferred)
  • * Experience using Generative AI (ChatGPT, GitHub Copilot, Amazon Q, Microsoft Copilot, or similar) to improve operational efficiency.
  • * Knowledge of AlOps platforms and Al-driven observability.
  • * Experience with Dynatrace Davis Al for anomaly detection and root cause analysis.
  • * Understanding of machine learning concepts for predictive monitoring and intelligent alerting.
  • * Experience building Al-assisted operational runbooks and troubleshooting workflows.
  • * Knowledge of Retrieval-Augmented Generation (RAG), vector databases, embeddings, and Al-powered knowledge search is a plus.
  • * Experience integrating Al with observability platforms using APIs.
  • * Familiarity with LLMs, prompt engineering, and Al-assisted automation.
  • * Experience using Python with Al frameworks (LangChain, LangGraph, OpenAI APls, or similar) is desirable.
  • * Exposure to Al-driven incident summarization, log analysis, and automated ticket enrichment.
  • Required Qualifications
  • * Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • * 6-10+ years of IT infrastructure or observability operations experience.
  • * 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch.
  • * Strong Linux system administration experience.
  • * Experience supporting enterprise-scale production environments.
  • * Strong troubleshooting and analytical skills.
  • * Excellent communication and stakeholder management Skills.
  • Preferred Certifications
  • * Dynatrace Associate or Professional Certification
  • * Splunk Enterprise Certified Administrator
  • * Elastic Certified Engineer
  • * Kubernetes (CKA/CKAD)
  • * AWS/Azure/GCP Certification
  • * ITIL Foundation
  • •AI/ML or Generative Al certification (preferred)
  • Soft Skills
  • * Strong ownership and accountability
  • * Excellent problem-solving and analytical thinking
  • * Ability to work independently with minimal supervision
  • * Strong collaboration across cross-functional teams
  • * Continuous learning mindset
  • * Ability to thrive in fast-paced production environments

Mention you found this on Data First Jobs — it helps us bring you more roles like this.

Machine Learning Engineer II- Premium

JPC TECHNO INC

Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.

Free, no spam. Unsubscribe anytime.