Job Description
We're looking for a senior DevOps Engineer to join a globally distributed team to help scale and operate infrastructure that powers millions of trades daily across CeFi and DeFi venues. You'll play a mission ciritcal role in ensuring performance, reliability and security of the systems in a demanding environment.
Key responsibilities:
- Ensure performance and reliability of mission-critical trading infrastructure by proactively identifying and resolving bottlenecks at all layers: compute, network, storage and application
- Support, operate, and continuously improve highly available, low-latency systems under global trading load, with a strong focus on automation and enabling self-service for internal teams.
- Participate in a day-time rotational on-call schedule, handling real-time operational incidents, root cause analysis, and long-term preventive improvements.
- Use metrics and observability tools (Prometheus, Grafana, etc.) to detect anomalies, monitor performance trends, and support incident response.
- Build, manage, and scale infrastructure-as-code using Terraform and configuration management tools like Ansible.
- Collaborate closely with trading, engineering, and security teams to align infrastructure improvements with application demands.
- Implement robust security practices, including IAM, threat detection, and secure CI/CD pipelines, to support high-trust production environments.
- Apply low-level optimization strategies to improve throughput, reduce latency, and increase system efficiency across diverse environments.
Interested in remote work opportunities in Devops? Discover Devops Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
Requirements:
- 5+ years in a DevOps/SRE roles focussed on high-performance, production-grade systems
- Proficiency in Linux admin and experience managing AWS cloud infrastructure
- Proficiency with containerised workloads (Kubernetes)
- Strong scripting and automation capabilities using Python and Bash; knowledge of Go, JavaScript, or TypeScript is a desireable
- Hands-on experience with observability stacks (Prometheus, Grafana) and production incident response
- Experience with automation and IaC tools; Ansible, Terraform
- Understanding of performance engineering at both the system and network levels, including tuning, profiling, and capacity planning
- Knowledge of cloud networking, security best practices, IAM and SIEM tools
- Experience working with crypto or finance is desireable
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
Similar Jobs
Explore other opportunities that match your interests