autor.

Buyer's guide

How to Evaluate an AI Development Studio in 2026

Ask every studio you are considering eight questions, and treat two of them as pass/fail: what is running in production right now with metrics you can see, and how they handle Canadian data privacy under PHIPA and PIPEDA. A studio that clears those two is worth a longer conversation.

How We Work

Why the bar moved

The barrier to calling yourself an AI development studio is zero. Hundreds of agencies stood up a website in 2023, built a few proof-of-concepts, and started charging enterprise rates. Meanwhile the technical bar for production AI went up: sub-second latency, PHIPA and PIPEDA compliance from the first line of code, systems that work at 2am on a Sunday and not just on demo day.

We have built over 50 AI products in four years, and we have been brought in to rescue projects other studios delivered — products that demoed well and then fell apart under real users. This is the evaluation we would run if we were hiring a studio ourselves, even if it was not us.

The 8 questions

Questions 1 and 4 are pass/fail. Question 5 is where most projects actually die — six months after delivery.

01

What is running in production right now, and can I see the metrics?

Not what a studio has built — what is handling real users and real edge cases today. When we answer this we show Loquent, our own voice AI platform: thousands of automated calls a month, an 89% automation rate, 4.3/5 patient satisfaction, sub-second first response — from live dashboards, not a pitch deck.

Red flag: We built AI for a big name but it is under NDA. Studios that ship production AI can describe architecture, scale, and outcomes without naming the client.

02

Walk me through a production incident and how you handled it.

Any studio with real production mileage has war stories. Ours: a silent vendor model update changed our transcription accuracy at 3am during a healthcare launch, because we had not pinned model versions. We published the full post-mortem. You want to hear what broke, how they detected it, time to fix, and what changed afterward.

Red flag: We resolved a minor issue quickly, with no detail behind it.

03

Who exactly will work on my project?

At Autor: senior engineers only, no offshore, no handoffs — the person in the sales call is the person writing your code. Many studios sell you an architect and deliver offshore contractors. For AI specifically, the gap between demo and production is a senior-engineering problem.

Red flag: We scale the team based on project needs.

04

How do you handle data privacy and compliance in Canada?

If you are in healthcare, dental, legal, or finance, this question eliminates most studios. We built Loquent with PHIPA and PIPEDA as architectural constraints: encrypted at rest and in transit, Canadian data residency, access-logged, full audit trail per patient interaction — reviewed by our clients' legal teams before we wrote code.

Red flag: We can add compliance later. Compliance is a design constraint, not a feature.

05

What happens after delivery? Who is on call?

Production AI is not build-and-forget: models drift, vendors push updates, volumes spike. Ask for SLA terms, monitoring specifics, and what happens Saturday at 3am. We run our own product in production 24/7, so monitoring is not a maintenance package we sell — it is how we already operate.

Red flag: A maintenance package with no response times attached.

06

Show me your testing process for AI-specific failures.

Unit tests do not catch prompt regressions, model drift, or hallucination on edge inputs. We replay 200 real call transcripts against every prompt change and test pinned vendor model versions against a saved corpus before promoting — a practice we adopted after a surprise model update dropped one city's automation rate nine points overnight.

Red flag: Testing an AI product the same way you would test a CRUD app.

07

What would you do differently if you started over?

A trap question, in the best way. We would build a graph-based conversation state machine instead of a linear flow, invest in observability from day one, and start our regression corpus at call one instead of month four.

Red flag: We would do everything the same. That studio has not run anything long enough to learn.

08

What would you tell us NOT to build?

The most valuable thing a partner does is talk you out of the wrong thing. We turned down a $200k project because the requirements would have produced something users did not need. Ask what a studio has declined.

Red flag: A studio that says yes to everything is optimizing for its revenue, not your outcome.

The evaluation checklist

Take this into every studio call. Six checks, and what a real answer sounds like.

CheckPass looks likeFail looks like
Production systemsLive metrics, current usersOnly completed handoffs
Canadian compliancePHIPA/PIPEDA by design, Canadian residencyWe will add it later, or HIPAA covers it
TeamNamed senior engineers, no handoffsWe scale the team as needed
Post-deliverySLAs, monitoring, on-callAn email checked in 48 hours
AI testingPrompt regression and model pinningWe have CI/CD
HonestyWar stories, declined projectsEverything went smoothly

Want the model behind our own answers? We wrote it down: senior engineers only, no offshore, no handoffs.

Frequently Asked Questions

Serious production builds typically start in the mid five figures. Autor bills $75–150/hr depending on scope, with a senior-only team and no blended rates, and Loquent — a full production voice AI platform — was built in under 8 weeks. Be suspicious of both extremes: $5,000 AI solutions and $500,000 discovery phases.

Evaluating studios right now?

Run these questions on us. We will tell you what we would build, what we would not, and whether we are the right fit. If we are not, we will say so.

hello@autor.ca

Autor Technologies Inc. · 401 Bay Street, Toronto · Founded 2021 · 50+ AI products delivered · 5M+ users impacted · $10M+ client funding raised