Back to Careers

Founding AI Quality & Release Engineer

Full-Time
U.S. Remote (Hybrid in NYC)

About Genesis Computing

Genesis Computing AI builds an agentic data engineering platform centered on Eve — an AI agent that autonomously executes complex data workflows inside enterprise cloud environments (AWS, Snowflake, Databricks, Azure). We're a seed-stage team of ~20, founded by early Snowflake employees and led by seasoned enterprise operators. Our clients include hedge funds, financial services firms, and large pharmaceuticals. We move fast, ship constantly, and build with the same AI-native tools we sell.

The Role

This is our first dedicated quality-infrastructure hire—and it is a software engineering role, not a traditional QA position.

You’ll write code most days: test harnesses, CI pipelines, release gates, and the LLM-agent systems that execute and evaluate our tests. You won’t spend your time manually running test cases. Agents do that here.

Your product is the verification and release system itself: the agentic test harness, evaluation logic, CI signal, release gates, deployment pipeline, and pre-production environments that carry every build from merge to customer.

Your mission: make a green build mean something, make upgrades safe, and make releases boring.

The role has two deliberately protected pillars:

  • AI-powered verification (~70%) — Own the trustworthiness of our test signal, including agent-run harnesses, LLM graders, regression coverage, and automated release gates. Close the gaps through which bugs currently escape to customers.
  • Release and environment engineering (~30%) — Own the mechanics of shipping, including the release pipeline, CI health, artifact publication, and the reliability of our demo and staging environments.

Make green mean green

  • Eliminate silent test skips across every lane. If a test cannot run, it must fail loudly—not quietly report green.
  • Harden our LLM-based graders so they are strict, calibrated, and adversarially reviewed. A lenient grader is a broken test.
  • Replace blanket retries with measured flake rates, a defined quarantine workflow, and a published flake dashboard.
  • Make the provenance of every result visible: what ran, what did not, why it failed, and whether the result can be trusted.

Build the release gates we’re missing

  • Automate fresh-install verification against real deployment targets, including Snowflake Native App installations with real secrets and external-access approval flows.
  • Build upgrade-from-previous-GA tests with explicit configuration-preservation assertions.
  • Verify that customer connections, permissions, model settings, and indexed data survive upgrades intact.
  • Turn our most damaging historical failure modes into mandatory, automated release gates.

Turn every escaped bug into a permanent test

  • Build the tooling and establish the policy: every customer-reported bug should produce an automated regression check, linked to its ticket, before the issue is closed.
  • Make regression coverage visible and auditable.
  • Drive re-reported bugs—issues that were fixed and later returned—to zero.

Own the agent-run UAT system as a product

  • Take ownership of approximately 26 natural-language UAT plans executed by coding agents in CI.
  • De-flake unreliable plans, refresh stale ones, and expand coverage as the product evolves.
  • Enforce the policy that feature changes include corresponding UAT updates.
  • Grow the release-gating suite from smoke-level coverage into a credible test of release readiness.
  • Decide which checks agents can own completely and which still require a small amount of high-judgment human review.
  • Remove the recurring release-verification burden from feature engineers.

Own the release pipeline end to end

  • Own release-candidate creation, cherry-pick workflows, and artifact publication across Docker Hub, Helm, Snowflake, and our managed cloud.
  • Handle the platform-specific constraints of shipping a Snowflake Native App, including security scans and version limits.
  • Replace tribal knowledge held by a rotating weekly “release master” with automation, clear runbooks, and a release status that nobody has to ask for.
  • Make releases repeatable, observable, and recoverable.

Keep CI and pre-production environments healthy

  • Own CI runner capacity and pipeline health.
  • Establish clear ownership and monitoring for API keys, service credentials, quotas, and certificates.
  • Automate certificate renewal and prevent expired credentials or exhausted quotas from unexpectedly breaking CI.
  • Maintain reliable internal demo and staging environments that sales engineers can trust during customer calls.

Improve production failure detection

  • Stand up error and exception tracking for customer deployments.
  • Connect production failures to actionable engineering context and regression coverage.
  • Help us detect failures before customers become our crash reporter.

What this role is not

It is not manual test execution

The manual workload exists because the automated signal is not yet trusted. Your job is to fix the underlying system—not staff the symptom.

It is not a release gatekeeper

You won’t be the person who relies on intuition to approve or reject releases. You’ll build systems that make release readiness visible and stop unsafe releases automatically, with evidence.

It is not a hidden 24/7 production on-call role

You’ll participate in business-hours escalations, as every senior engineer does, and you’ll make incidents rarer and easier to diagnose. Formal after-hours production on-call is a separate function that we will staff as the need matures.

About you

  • You are a strong software engineer who has built test, CI, release, or developer infrastructure. Your background might be in platform engineering, developer productivity, quality engineering, SDET-adjacent engineering, or release engineering.
  • You write production-quality Python. TypeScript experience is a plus.
  • You are deeply fluent in GitHub Actions or an equivalent CI system and are comfortable owning complex pipelines end to end.
  • You’re genuinely excited about LLM- and agent-driven verification. You want to build agent runners, design and calibrate LLM graders, improve prompts and evaluation rubrics, and make probabilistic systems produce trustworthy release signals.
  • You understand that an LLM grader is itself software that must be tested, versioned, measured, and challenged.
  • You can work effectively with containers, Kubernetes, cloud deployment mechanics, secrets, credentials, and certificates. You don’t need to be a career SRE, but infrastructure cannot intimidate you.
  • You think like an owner. You measure success by escaped defects and release confidence—not by the number of bugs filed.
  • You would rather eliminate the cause of a flaky test than add another retry.
  • You are senior enough to challenge a release when the evidence says it is unsafe—and pragmatic enough to build a better system so the same judgment does not depend on you next time.
  • You communicate clearly. Runbooks, retrospectives, dashboards, and operating policies are part of how your work scales beyond one person.

Nice to have

  • Snowflake Native Apps or Snowpark Container Services experience
  • Playwright or similar browser-automation experience
  • LLM evaluation, LLM-as-judge, or grader-design experience
  • Experience testing agentic or other nondeterministic systems
  • Kubernetes and Helm experience
  • Experience building quality or release infrastructure at an early-stage startup

How we’ll measure success

During your first 12 months, we expect to see:

  • Post-GA regressions per release: reduced from approximately one per release to near zero
  • Release candidates per release: reduced from the current 9–15 to the low single digits
  • Manual UAT cases per engineer per week: reduced from 26–36 to fewer than 10
  • Bug discovery: shifted from customer-first to internal-first
  • Re-reported regressions: zero
  • Silent test skips: zero
  • Flake health: measured, owned, and visible through a trusted dashboard
  • Release readiness: understandable from the system’s output without relying on tribal knowledge or asking for a status update

Ultimately, success means engineers can move faster because they trust the verification system, customers encounter fewer regressions, upgrades preserve everything they should, and releases become routine rather than heroic.

Our Development Philosophy

  • Every engineer is now "at minimum a technical manager" — managing AI agents to do the work, not hand-coding
  • Higher-level thinking is required — product mindset, architecture, communication (to agents and humans)
  • Independent delivery — no spoon-feeding, hit the ground running
  • Everyone is Multidisciplinary — one day CI/CD, next day AWS infra, next day frontend dashboard, next day LLM expert

Why Join Us?

  • Forefront of software development's transformation, seasoned leadership (ex-Snowflake founders, ex-Goldman head of engineering), ownership and speed, not a waiting-for-tickets culture
  • Shape everything — as an early team member, you influence product direction, technical architecture, and engineering culture
  • Real enterprise impact — we build end-to-end solutions for demanding customers, not toy demos
  • Cutting-edge stack — AI agents, multi-agent orchestration, data automation, and cloud-native systems
  • Competitive compensation — salary, equity, and benefits
  • Flexible environment — remote-first with periodic offsites, NYC-office vibes for those in the area

How to Apply

Ready to build something extraordinary? Send your resume, a link to your portfolio or GitHub, and a brief note about why you’re excited to join Genesis Computing to careers@genesiscomputing.com. Let’s create the future together.