Kobold Help

QA checks explained

What Moss tests on the preview theme before a change reaches you.

Moss tests the copy of your theme on your real storefront, at 390px (phone) and 1440px (desktop), on the pages where the change shows plus any pages you added under Agent rules → QA pages. Every check streams into the QA card as it finishes. Checks run side by side (pages and widths in parallel, the scripted steps and smoke tests alongside them), so a typical run takes about two minutes.

The checks are written for this task. From the plan and what Nib changed, Moss picks the pages to test, the steps that exercise the change (open it, use it, the empty and error states) and the elements the change is meant to alter. A task that adds a cart drawer field is tested with the drawer open, not with a generic walk through the store.

GroupWhat Moss checksPasses when
StaticTheme Check (Shopify's linter), valid JSON templates and locale files, every new translation key exists in every store language, no new render-blocking scripts0 new Theme Check errors. Errors your live theme already has never count, even in a file Nib edited. Warnings are reported but don't block
VisualA visual diff of live vs preview for each page and width, with animations offNothing changed outside the area the plan targets
ScriptedSteps written for this task, e.g. open a product that's in stock, add to cart, open the drawer, type WELCOME10, check the totalEvery step passes
SmokeAdd to cart, change quantity, remove, go to checkout (never submits), and your app blocks still renderAll pass
Console & speedNo new console errors (Shopify's own analytics scripts are ignored), and a Lighthouse performance score on mobile and desktopPreview scores at least the live score minus 5

Moss may ask for test data

Some changes can only be tested with things only you know: a discount code that works and one that doesn't, a customer tag, a product that's sold out. Pip asks for these while planning; if something is still missing when testing starts, Moss asks in the thread with the same question card (one question at a time, your own words, or skip). QA waits for your answers, then starts. You can also just type a reply; it counts as your answer.

The QA report

Every QA run is its own card in the thread, so a fix round adds a new card below the earlier one instead of rewriting it. Open QA report on a card shows that run in the side panel:

  • Moss's summary, run time, credits and retries;
  • Theme Check errors with file and line, JSON and locale problems;
  • the visual diff for every page and width, with the changed pixels in red and links to the before and after captures;
  • each scripted step with its screenshot, and why a step failed;
  • console errors and Lighthouse scores (phone and desktop, copy vs live).

When a task has several rounds, pick one at the top of the report.

Scripted steps wait for your store after each click or keypress: Moss holds the next step until Shopify's cart requests (add, change, update) have finished, so "open the cart" after "add to cart" sees the item.

Lighthouse doesn't hold up your review: the change is ready as soon as the other checks pass, and the Lighthouse row fills in a moment later. If the copy scores lower than your live theme by more than your tolerance, Moss adds a note in the thread. Your live theme's score is remembered until the live theme changes, so most runs only test the copy. Lighthouse has a four-minute limit. If it runs longer or can't start, that check is marked skipped and the rest of QA carries on. If a QA run ever stops making progress (for example after a restart on our side), Moss picks it up again within a few minutes.

Masking

Moss ignores the area the plan is meant to change, anything that moves on its own (carousels, countdowns, lazy images, detected by capturing live twice) and any element you mark with data-kobold-ignore.

Before and after

Moss replays how a shopper reaches the change (open the product, add to cart, the drawer opens) on your live theme and on the copy, and captures the screen at both widths. Those are the Before / After shots in the review panel, so you see the feature itself rather than an untouched page. They show up as soon as Moss has run, even before every check passes.

Shopify doesn't let a storefront load inside another site, so the Preview tab shows Moss's capture of the copy. Use Open the live preview to click around the copy in a new tab.

When a check fails

A failed check never sends the whole QA round again, and never goes back to Nib on its own:

  1. Flukes are retested alone, once. A timeout, a slow page or a network error re-runs only that check. A wrong value or a visual change isn't a fluke, so it goes straight to Moss.
  2. Moss reads what still fails. He can open the theme files involved to find the cause, so his suggested fix names the file and the change. He compares it with the plan and what Nib changed, and works out why. He can accept a check that only flags the intended change or something that isn't this task's doing, like Shopify's own scripts. He only accepts visual, console and speed checks, never a broken step or a failing smoke test, and the QA card marks those rows "Accepted by Moss" with his reason under the row; the title counts them separately ("15 passed · 2 accepted").
  3. You decide what's next. If something is still failing, the task moves to Needs your help with Moss's explanation and three options: Ask Nib to fix it (with the fix Moss suggests), Send it for approval (skips those checks without testing again, so you review it yourself) or Talk to Nib.

Nothing is published until you approve it.

If your store blocks automated visitors, see Allow-list the KoboldQA bot.

On this page