Site Reliability Engineer – Data Infrastructure & Distributed Systems
Westbury Partners Sydney, AustralienSite Reliability Engineer – Data Infrastructure & Distributed Systems
Own and evolve large-scale data infrastructure, combining SRE practices, Linux expertise, distributed systems, automation, and operational ownership to keep critical platforms reliable, scalable, and high-performing.
What You'll Do:
Join an experienced infrastructure team responsible for the reliability, scalability, and evolution of large-scale data platforms. You’ll operate complex distributed systems, solve production challenges, automate infrastructure, and work directly with technical users who need fast, reliable solutions.
Your responsibilities will include:
- Operate and monitor distributed data platforms, including Kafka, HDFS, Dremio, Kubernetes, and in-house data pipelines.
- Manage deployments, upgrades, capacity planning, infrastructure changes, and failure domains across large-scale environments.
- Build automation and CI/CD solutions for rapid, reliable deployment of hardware and software.
- Participate in on-call rotations, investigate incidents, improve alerting, and implement preventative engineering.
- Diagnose Linux, networking, storage, memory, filesystem, and performance issues across self-managed infrastructure.
- Partner directly with traders, researchers, developers, and engineering teams to troubleshoot problems and design effective data solutions.
- Evaluate emerging technologies and implement improvements that strengthen reliability, scalability, and operational efficiency.
- Develop Python-based infrastructure tooling, health checks, automation, and self-service solutions.
- Contribute to cross-office engineering initiatives and collaborate with teams across multiple international locations.
Why Join Us:
- Work with genuinely large-scale, multi-petabyte data infrastructure.
- Take ownership of systems where reliability, performance, and engineering judgment matter.
- Learn from an experienced infrastructure team with a structured development path.
- Gain hands-on exposure to complex open-source technologies and distributed systems.
- Build engineering depth across Linux, networking, Python, Kubernetes, Kafka, HDFS, automation, and system design.
- Make improvements that directly affect researchers, developers, and other highly technical users.
- Grow from strong fundamentals rather than being expected to know every part of the data ecosystem on day one.
About You:
You’re an infrastructure-minded engineer who enjoys understanding how systems behave under pressure and solving problems at their source.
You’ll ideally bring:
- Hands-on production SRE or infrastructure experience, including meaningful on-call and incident ownership.
- Experience operating self-managed infrastructure such as bare metal, datacentre environments, VMs, or Kubernetes.
- Strong Linux fundamentals, including processes, filesystems, networking, memory, storage, and performance troubleshooting.
- Genuine operator-side experience with at least one of Kafka, HDFS, or Kubernetes.
- Experience with infrastructure automation, configuration management, CI/CD, or deployment tooling.
- Python experience focused on infrastructure, automation, health checks, or operational tooling.
- A proactive approach to reliability, capacity, failure modes, observability, and continuous improvement.
- The ability to explain what you personally delivered, why it mattered, and how you measured its impact.
- Curiosity about complex technology and a willingness to develop deeper expertise in unfamiliar systems.
You don’t need to know everything on day one. You do need to know how to operate, investigate, automate, and keep learning. If you’re ready to go deep into challenging infrastructure and build systems that people can depend on, this is an opportunity to make a real impact.