Skip to content
Andres Carreon

NowQuality & AI automation at a fintech startup

I break software on purpose.

I'm Andres Carreon, a QA engineer who writes a lot of code. For nine years I've tested iOS, Android, desktop and web apps so they ship without surprises. After hours I build my own tools and write down what I find.

A recorded jevvium run. Goal: A registered user logs in with a valid email and password and is told they are logged in. Four decisions: tap Login, type the email and password, tap LOGIN, goal reached. Passed in 5.2 seconds, and the replay with plain Appium passed too.

Real run · jev-1.13.0 · iOS 26.5 simulatorSource
9 years
9yearstesting software, from enterprise to crypto to fintech
1K+ tests
1K+testsautomated across the suites I built from zero, on mobile, desktop and web
40K minutes
40Kminutesof CI saved every month with self-hosted runners I set up
2 products
2productsof my own with paying customers

Now building · open source

jevvium turns acceptance criteria into Appium tests.

You describe what a feature should do. A decision model finds the path through the app one tap at a time, and jevvium writes it down as a plain WebdriverIO test that CI runs with no AI in the loop.

Read the code on GitHub

1Describe it

A goal, some test data and what should be on screen when it works. A few lines of YAML.

goal: A registered user logs in…
inputs:
email: qa.demo@example.com
expect:
- text: You are logged in!

2Explore once

A decision model picks the next action from elements that really exist on screen, and stops to escalate when it isn't sure.

  • tap "Home"
  • tap "Webview"
  • tap "Login"chosen · 0.87
  • tap "Forms"
  • tap "Swipe"
  • tap "Drag"

3Run forever

The path becomes a normal test. jevvium replays it once with plain Appium to prove it passes, then CI owns it.

$ npm run test:generated:ios
✓ login
✓ login-invalid-email
✓ forms-switch
no model calls at run time

47.7s → 16.3s

to explore four criteria, after rebuilding how it reads, decides and taps

~25 ms

per simulator tap through a small native helper, instead of about 400 ms

< $0.001

of model cost for those four criteria

4 at once

simulators exploring in parallel, one Appium server each

Field note: it found a real bug

In Sauce Labs' demo shop, “price, low to high” puts $7.99 after $49.99.

The app compares prices as text. The model couldn't see prices, so it wasn't sure the goal was reached. The expect check is what caught it.

Day job

Nine years of making software behave.

Quality engineering at a global consultancy, a crypto wallet used by millions, and now a fintech startup. Here it is as release notes.

  1. v3currentAug 2026 to now

    Fintech startup

    Quality & AI automation engineer

    name under NDA

    • Set up self-hosted GitHub Actions runners for Linux and macOS builds, saving about 40,000 CI minutes a month.
    • Built the eval suite and knowledge base for a customer-support AI agent, now at a 94% pass rate.
    • Ship production code beyond QA: KYC backend validations, a React Native migration and bug fixes.
  2. v2May 2022 to Aug 2026

    Crypto wallet

    SDET, quality engineering

    self-custody, used by millions

    • Built the end-to-end automation from zero: about 600 tests across mobile, desktop and browser extension, running twice a day.
    • Each run replaced about five hours of manual regression per platform.
    • Built an AI bot that reads each pull request, decides what to test and opens a PR with the new tests.
    • Slack run reports with AI triage that tells a product bug from an environment problem.
    • Led four QA engineers and owned release sign-off.
  3. v12017 to May 2022

    Accenture

    QA engineer, test automation

    enterprise systems

    • Led a QA team of 15 building automated regression suites and page-object frameworks in CI.

After hours

Products I built on my own time.

Nine shipped so far, two with paying customers. Building them is where most of my AI tooling comes from, including the agent that triages my own bugs.

Plus two open-source test frameworks, a hackathon soap opera and a few experiments.

All projects

Writing

Notes worth keeping.

Things I found interesting enough to write down: AI, attention, and building things alone. Most of them in English and Spanish.