JPC TECHNO INC
Machine Learning Engineer II- Premium
Contract · In Office · Phoenix, Arizona (USA)
Posted Aug 4, 2026
Work Options
Industry
Job Type
Positions
Position Group
- We are seeking a highly skilled Senior Observability Operations Engineer to manage and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kuberetes, Linux, and cloud-native observability solutions.
- Experience leveraging AI/ML and Generative Al to improve observability, automate operations, and accelerate incident resolution is highly desirable.
- The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications.
- Key Responsibilities
- * Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
- * Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
- * Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
- * Configure Dynatrace OneAgent, ActiveGate, Synthefic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis Al, and Application Performance Monitoring (APM).
- * Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
- * Develop dashboards; alerts, reports, and executive operational metrics.
- * Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
- * Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
- * Perform root cause analysis for production incidents using observability platforms.
- * Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
- * Automate operational tasks using Python, Shell scripting, REST APls, Terraform, or Ansible.
- * Participate in incident, problem, change, and release management processes.
- * Drive platform upgrades, patching, security compliance, and operational governance.
- * Improve platform reliability through automation, self-healing, and Al-assisted operations.
- Required Technical Skills
- Observability Platforms
- * Dynatrace Administration
- * Splunk Enterprise Administration
- * OpenSearch Administration
- * Elasticsearch Administration
- * Grafana
- * Prometheus
- * Kibana
- * Jaeger
- * Open Telemetry
- * Kafka (preferred)
- Infrastructure
- * Linux Administration
- * Kubernetes
- * Docker
- * OpenShift or Rancher
- * Networking (TCP/IP, DNS, Load Balancers, Firewalls)
- * System Administration
- Cloud & DevOps
- * AWS, Azure, or GCP
- * CI/CD pipelines
- * Git
- * Terraform
- * Ansible
- * REST APIs
- Scripting
- * Python
- * Bash/Shell
- * PowerShell (preferred)
- AI & Automation Skills (Preferred)
- * Experience using Generative AI (ChatGPT, GitHub Copilot, Amazon Q, Microsoft Copilot, or similar) to improve operational efficiency.
- * Knowledge of AlOps platforms and Al-driven observability.
- * Experience with Dynatrace Davis Al for anomaly detection and root cause analysis.
- * Understanding of machine learning concepts for predictive monitoring and intelligent alerting.
- * Experience building Al-assisted operational runbooks and troubleshooting workflows.
- * Knowledge of Retrieval-Augmented Generation (RAG), vector databases, embeddings, and Al-powered knowledge search is a plus.
- * Experience integrating Al with observability platforms using APIs.
- * Familiarity with LLMs, prompt engineering, and Al-assisted automation.
- * Experience using Python with Al frameworks (LangChain, LangGraph, OpenAI APls, or similar) is desirable.
- * Exposure to Al-driven incident summarization, log analysis, and automated ticket enrichment.
- Required Qualifications
- * Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- * 6-10+ years of IT infrastructure or observability operations experience.
- * 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch.
- * Strong Linux system administration experience.
- * Experience supporting enterprise-scale production environments.
- * Strong troubleshooting and analytical skills.
- * Excellent communication and stakeholder management Skills.
- Preferred Certifications
- * Dynatrace Associate or Professional Certification
- * Splunk Enterprise Certified Administrator
- * Elastic Certified Engineer
- * Kubernetes (CKA/CKAD)
- * AWS/Azure/GCP Certification
- * ITIL Foundation
- •AI/ML or Generative Al certification (preferred)
- Soft Skills
- * Strong ownership and accountability
- * Excellent problem-solving and analytical thinking
- * Ability to work independently with minimal supervision
- * Strong collaboration across cross-functional teams
- * Continuous learning mindset
- * Ability to thrive in fast-paced production environments
Mention you found this on Data First Jobs — it helps us bring you more roles like this.
Machine Learning Engineer II- Premium
JPC TECHNO INC
Similar Engineering Jobs
View all Engineering jobs→Dotmatics
Senior Data Analytics Engineer
New
RemoteRemote
Tata Consultancy Services
Power BI Developer
New
Toronto, Ontario (Canada)$100,000 - $120,000
Veolia | North America
Senior Mechanical Commissioning Engineer (Data Center)
New
Richmond, Virginia (USA)
Digital Biz Tech
BI Developer
New
Portsmouth, New Hampshire (USA)
Intuition IT – Intuitive Technology Recruitment
Data Engineer
New
Raleigh, North Carolina (USA)
Jobright.ai
Machine Learning Engineer, Applied AI — New Grad
New
Canada
Like this role? Get carefully selected jobs like it, twice a week, straight to your inbox.
Free, no spam. Unsubscribe anytime.