Current opportunity
Site Reliability Engineer
Oxio
About the opportunity
Site Reliability Engineer OXIO is the first NeoTelco. We are building the world’s largest, most accessible, and insightful Telecom network. Our platform empowers anyone to spin up their own carrier from a browser, scaling and supporting you as you scale your network to millions of users. We ensure that users and devices are connected, and stay connected wherever they go: Cross- country, carrier, or cellular technology.
We help them pay less for mobile data. This technology is provided through our Carrier-as-a-Service platform: BrandVNO, a fully customizable telecom service. In addition, we enable clients of our service to extract the value from telecom data - enriching their customer experience, business intelligence, and product understanding in the many markets in which we operate.
Come join us in creating a modern technology platform with a group of engineers dedicated to advancing our vision. Our team is passionate about what we build, open to new ideas and challenges, and has our sights set on the future of connectivity.
Responsibilities
- Design and implement platform on the cloud to support OXIO backend servicesAutomate technical operations: deployments, scaling, recovery, etc.
- Monitor and maintain mission-critical production infrastructure to ensure maximum uptimeParticipate in an on-call rotation and culture of continuous improvement through blameless postmortemsEnable the Engineering/Telecom/Data Engineering teams by providing them the tools to operate the service they build EssentialsUnderstanding of Linux/Unix systems (most systems are Linux-based).
- Familiarity with Linux/Unix system internals like process management, filesystems, memory management, and networking.
- Proficiency in at least one programming language (Python, Go, or Ruby) and strong skills in scripting (Bash, Perl).
- Experience with infrastructure provisioning tools such as Terraform, CloudFormation, or Ansible.
- Familiarity with containerization (Docker) and orchestration tools (Kubernetes).
- Familiarity with monitoring tools like Prometheus, Grafana, or Datadog.
- Knowledge of setting up alerts, analyzing logs, and creating dashboards for observability.
- Familiarity with incident management practices (e.g., runbooks, postmortems).
- Experience in being part of an on-call rotation and handling incidents.
- Experience in setting up and maintaining Continuous Integration/Continuous Delivery pipelines (Jenkins, GitLab CI, CircleCI, etc.).
- Hands-on experience with cloud providers (AWS, Google Cloud, Azure).
