Shopify Storefront Agent Evaluator
$40-$50 / hr
$40-$50
About the role
Shopify's Storefront Agent is the AI shopping assistant that sits on a merchant's store — it searches the catalog in natural language, recommends products, builds carts, and answers questions about shipping, returns, and store policies.
Your job is to stress-test it and grade what comes back. You'll send the agent a handful of realistic shopper prompts, then score each response against a short rubric. The work is straightforward and self-contained: no coding, no data pipelines, no long-form writing.
What you'll do
-
Prompt the Storefront Agent as a real shopper would — product discovery, comparisons, sizing and stock questions, shipping and returns, edge cases and awkward requests
-
Grade each response on a small set of dimensions:
-
Relevance — did it actually answer what was asked, within the constraints the shopper gave?
-
Accuracy — is every claim (price, stock, policy, delivery window) true and grounded in the store's own catalog and policy data, rather than plausible-sounding invention?
-
Safety — did it avoid harmful advice, unsupported claims, and actions it wasn't authorized to take?
-
-
Leave a short, concrete note on anything that failed, so the reason is legible to someone who wasn't there
-
Commitment: up to 5 hours/week
Qualifications
Required:
-
Consumer e-commerce fluency — you shop online regularly and have a clear sense of what a good vs. a useless product recommendation looks like.
-
Careful, consistent judgment — you can apply the same rubric the same way across many responses, and separate "this response was wrong" from "this response wasn't what I'd have written."
-
Clear, concise written English — enough to explain in a sentence or two why a response failed.
Strongly preferred:
-
Hands-on Shopify merchant experience — you've run or operated a Shopify store and know how shoppers actually behave on one.
-
Experience using AI shopping assistants or chat agents as a customer, and a feel for where they tend to break.
Preferred (nice to have — we'll ramp you on the specifics):
-
Prior work evaluating, red-teaming, or annotating AI model outputs.
-
Familiarity with retail operations: catalog and variant structure, inventory, shipping and returns policy.
What this is not
This is not a customer-support role and not an engineering role. You are not fixing the agent — you are judging it, precisely and repeatably, so the team can measure where it falls short.
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
Contract and Payment Terms
- You will be engaged as an independent contractor.
- This is a fully remote role that can be completed on your own schedule.
- Projects can be extended, shortened, or concluded early depending on needs and performance.
- Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution.
- Payments are weekly on Stripe or Wise based on services rendered.
- Please note: We are unable to support H1-B or STEM OPT candidates at this time.
About Mercor
Mercor partners with leading AI labs and enterprises to train frontier models using human expertise. You will work on projects that focus on training and enhancing AI systems. You will be paid competitively, collaborate with leading researchers, and help shape the next generation of AI systems in your area of expertise.
Earn up to $200 by referring
Posted 7 hours ago