Open position
Lead SRE (Site Reliability Engineer)
Global Remote ·Full-time
About SQD
At SQD, we're redefining the database layer for the AI and Web3 world. Our flagship product, Portal, streams validated real-time and historical data from 190+ networks to developers building across DeFi, AI agents, and Web3 ecosystems. As we scale, we need an experienced SRE to help take our systems to enterprise-grade infrastructure.
What you'll do
- Help design, build, and optimize a high-availability blockchain data ingestion pipeline with sub-second latency and 99.9% uptime.
- Own our CI/CD, orchestration, and infrastructure-as-code layers.
- Identify and implement the right tools to run and monitor blockchain nodes, with failover solutions using multiple node providers.
- Continuously assess infrastructure trade-offs (hosted nodes, subscriptions, bare metal) to achieve optimal performance, reliability, and cost-efficiency.
- Build and maintain a public status page, and the incident management process behind it: severities, escalation paths, and post-mortems that change how we build.
- Work closely with engineers integrating new chains, making necessary patches to the ingestion pipeline.
- Define and maintain key SRE metrics, logging, and alerting, proactively identifying and resolving reliability risks.
- Stay ahead of past incidents, continuously improving observability, automation, and fault tolerance.
- Contribute to incident response, troubleshooting, and on-call rotations.
Requirements
- 3+ years of experience as an SRE, DevOps Engineer, or similar role, with reliability achievements you can point to and talk through.
- Experience running production services against real SLAs, including on-call, incident response, and post-mortems.
- Experience defining and implementing metrics, logging, and alerting to keep production healthy and prevent incidents.
- Proficiency in Kubernetes, Terraform, Prometheus, Grafana, or equivalent monitoring tools.
- Strong understanding of distributed systems, streaming data pipelines, and the failure modes that come with them.
- Deep knowledge of cloud infrastructure (AWS, GCP, or bare metal setups) and cost optimization strategies.
- Ability to balance performance, reliability, and cost, assessing when to use hosted nodes, subscriptions, or self-hosted setups.
- Programming skills (Python, Go, Rust, or Bash) for automation and infrastructure tooling.
- Experience monitoring and running blockchain nodes (Ethereum, Solana, etc.), and/or working with node providers as a backup solution.
- Willingness to learn and understand the internals of blockchain nodes, EVM/SVM data, and work with engineers to integrate new chains.
- Previous experience in Web3 is preferred.
Benefits
- Competitive salary + token incentives
- Fully remote with flexible hours
- High-impact role with ownership where your work directly shapes the reliability of the onchain data layer
- Build the operational foundations for a frontier AI/Web3 company
Get in touch
Email c.cliff@sqd.dev or message @connorcliff on Telegram.