"system observability" Jobs
3867 open tech roles matching “system observability”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AWS. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 3867 results
Join Okta as a Staff Software Reliability Engineer to design and build scalable data platform services in a hybrid work environment.
Join a small, independent team at Cloudinary to build a scalable backend for a visual moderation platform.
Join Datadog as a Senior Software Engineer to enhance a petabyte-scale distributed storage system for the AI era.
Join Mind Robotics as a DevOps Engineer to build and operate infrastructure for advanced robotic systems in a hands-on AI environment.
Join Composio as an SDET to build automation frameworks and improve quality practices for integrations and AI-powered systems.
Design and implement low-level systems software for GPU clusters in a pioneering AI infrastructure company.
Join TurbineOne as a Senior/Staff Software Engineer in DevOps, focusing on building automated testing infrastructure for edge hardware.
Join Legora as a Senior Platform Engineer to enhance infrastructure reliability and performance for a fast-growing legal tech company.
Build and operate cloud infrastructure for mission-critical aviation systems in a hands-on role focused on reliability and performance.
Join Anthropic as a Staff+ Software Engineer to build reliable connectivity infrastructure for AI systems.
Join Arena as a Software Engineer focusing on product security, designing and implementing robust security systems for AI applications.
Join MongoDB as a Senior Platform Engineer to enhance developer productivity through a self-service internal development platform.
Own the reliability and security of fal's generative media model APIs in a hybrid ML Engineering/SRE role.
Join Valon Labs as a Senior Product Manager to drive the roadmap for core infrastructure domains in a high-growth environment.
Design and maintain data infrastructure for AI systems, ensuring reliable performance in sensitive environments.