Descripción
Wrike is the most powerful work management platform. Built for teams and organizations looking to collaborate, create, and exceed every day, Wrike brings everyone and all work into a single place to remove complexity, increase productivity, and free people up to focus on their most purposeful work. Our vision: A world where everyone is free to focus on their most purposeful work, together. About the Role: Wrike’s Backend Reliability (BRE) team is the backbone of our backend infrastructure and the guardian of our uptime. Our mission is to achieve and sustain 99.99% availability while building the tools, components, and safety nets that the entire engineering organization relies on. As a Senior / Staff Backend Engineer on this team, you won’t just close tickets - you’ll architect core reliability solutions that shape how Wrike scales, performs, and recovers from failure. Your Impact: Design, build, and maintain critical reliability components such as HTTP rate limiters, internal DB schema migration tools, circuit breakers, and distributed Redis-based caching. Troubleshoot complex production issues, optimize PostgreSQL usage, and ensure our distributed systems remain performant and stable under high load. Lead preliminary investigations during severe production incidents: identify likely root causes, assess impact, and propose mitigation options. The long-term fixes are then implemented by the owning team, based on your findings. Create scalable, reusable tools and frameworks that help other engineering teams build more resilient services. Leverage AI-powered tools and coding agents to accelerate development, analyze architectures, and automate repetitive or error-prone tasks. Influence reliability best practices across engineering by sharing knowledge, reviewing designs, and setting high technical standards. Your Qualifications: Strong expertise with Java/JVM, building scalable, high-performance backend systems; open to leveraging other languages when appropriate. Solid understanding of distributed systems concepts, including high availability, CAP theorem, and fault tolerance. Deep experience with relational databases (PostgreSQL) and key–value / non-relational s
torages (Redis). Practical experience with containerization and cloud-native environments, including Docker and Kubernetes. Hands-on experience with message brokers such as RabbitMQ or Kafka.
Postular ahora
Publicado 30/7/2026