Job detail for QA Engineer (AI & Automation)
C
QA Engineer (AI & Automation)
Cyber Nest
Todayvia fourdayweek
Use AI to assess how you fit
About the Role
We are looking for a mid-level QA Engineer to own quality for a platform that integrates with multiple AI providers and uses them for data enrichment. You will combine manual testing, test automation, and AI-specific validation to make sure enriched outputs are accurate, consistent, and reliable in production.
****
- Design and execute manual test plans, test cases, and exploratory testing for web and API features.
- Build and maintain automated test suites for UI, API, and integration layers, and wire them into CI/CD.
- Validate AI-enriched outputs for accuracy, relevance, consistency, and hallucination risk, using golden datasets and evaluation criteria.
- Test integrations with multiple AI providers (such as OpenAI, Anthropic, and others), including failure handling, timeouts, rate limits, fallback behavior, and provider response variations.
- Compare output quality across providers and model versions, and flag regressions when models or prompts change.
- Test prompt changes, edge cases, adversarial inputs, and data quality issues in the enrichment pipeline.
- Monitor cost, latency, and token usage behavior in test scenarios and report anomalies.
- Log clear, reproducible bugs and work closely with developers, product owners, and DevOps to resolve them.
- Contribute to QA processes, documentation, and release sign-off.
Required Skills
- 3-5 years of QA experience covering both manual and automation testing.
- Hands-on experience with automation tools such as Playwright, Cypress, or Selenium.
- Strong API testing skills (Postman, REST Assured, or pytest/requests-based frameworks).
- Solid understanding of testing non-deterministic systems, including LLM output evaluation, prompt testing, and handling variability in results.
- Experience testing third-party integrations, webhooks, and asynchronous or queue-based workflows.
- Familiarity with CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins) and version control with Git.
- Working knowledge of SQL and basic scripting in Python or JavaScript/TypeScript.
- Strong written communication and bug reporting skills.
Nice to Have
- Experience with AI evaluation tools or frameworks (such as promptfoo, DeepEval, or Ragas).
- Understanding of embeddings, RAG pipelines, and vector databases.
- Experience with performance and load testing (k6, JMeter, Locust).
- Exposure to Rails or Python/FastAPI backends, Redis, and background job systems like Sidekiq.
- Familiarity with AWS environments and log/monitoring tools.
- Experience with data-heavy or healthcare-adjacent products, including data privacy awareness.
What Success Looks Like
Within three months, you will have a stable automated regression suite covering core enrichment flows, a repeatable AI output evaluation process, and clear quality metrics that the team trusts for release decisions.
What We Offer
- Work on a real production AI platform with multiple provider integrations.
- A collaborative engineering team and room to shape QA practices.
- Competitive salary and growth opportunities.