BPO Service
LLM evaluation services: human review and AI agent testing
Calibrated reviewers score your model's outputs against your rubric and test your agents on real tasks.
Published · Last updated
Direct answer
Actigy BPO provides LLM evaluation services as a nearshore business process outsourcing (BPO) company headquartered in Prague, Czech Republic. Reviewers score large language model (LLM) outputs against client rubrics and check agent tasks in approved tools. A process audit defines the sample before a paid pilot. The client keeps model and release decisions.
Your rubric · Approved test access · Client-owned release decisions
Evidence for your model team
Actigy BPO records scores, examples and unresolved questions. Your team decides whether the evidence supports a change or release.
Key takeaways
- Every Actigy BPO engagement starts with a process audit and a paid pilot.
- Actigy BPO works in the client's tools, with access limited to the systems the client approves.
- Compare scores by task, sample and rubric version.
- Keep model changes and release approval with the client.
LLM evaluation services at a glance
Actigy BPO scopes human review around the client's task set, tools and decision owners.
| Service | Human scoring, response comparisons and agent task review |
|---|---|
| Who it is for | Artificial intelligence (AI) teams, model builders and chatbot owners |
| Delivery locations | Bulgaria, Romania, Poland and Ukraine |
| Coverage | Business hours to 24/7, set per engagement |
| Pricing model | Per full-time equivalent (FTE) by role; written quote after the process audit |
| Typical start | Pilot usually 2 to 4 weeks after the process audit, subject to readiness |
Human evaluation for LLM outputs
Actigy BPO reviews whether an answer meets the client's task and rubric. Evaluation checks artificial intelligence (AI) outputs before and after changes. Automated scores help compare repeatable tests. Human review adds context, reasons and examples that a single score can hide.
| Method | Human review in the client's workflow |
|---|---|
| Automated metrics | Examine flagged examples; the client owns metric code and thresholds. |
| Model-graded evaluation | Check a sample of model-assigned scores against the approved rubric. |
| Human rubric scoring | Score answers and record evidence for each judgment. |
| A/B and side-by-side tests | Compare responses under the same conditions; record preferences and reasons. |
| Safety and policy checks | Flag breaches of the client's rules; escalate unclear policy cases. |
How Actigy BPO runs human evaluation for LLMs
Actigy BPO separates tasks so the report shows what each score means. A factuality check compares an answer with supplied evidence. A preference judgment compares responses. Neither task gives reviewers authority to change the model or its policy.
| Task | Review record |
|---|---|
| Rubric scoring | Score, criterion and supporting example |
| Response comparison | Preferred response, tie or unresolved choice, with a reason |
| Factuality and grounding | Claim checked against the supplied source, including missing support |
| Tone and brand | Match or conflict with approved style examples |
| Safety and policy | Policy category, evidence and escalation owner |
| Agent tasks | Task outcome, tool step and handoff record |
Calibration, pilot and review method
Actigy BPO agrees the guideline, reference set and review responsibilities before production. Quality assurance (QA) measures consistency against that agreement. Your team approves changes to the rubric.
- Process audit. Map tasks, sample limits, tools, access and decision owners.
- Guideline design. Write standard operating procedures (SOPs), score definitions and escalation rules with the client.
- Reviewer selection. Check task understanding against approved examples before assigning live work.
- Calibration. Compare reviewers on gold items with agreed reference answers; resolve differences.
- Paid pilot. Test real examples, unclear cases, rework and report fields.
- Scale gate. Add work only after the client accepts the pilot against agreed thresholds.
- Continuous review. Sample work, examine defects and approve guideline updates before the next batch.
Evaluation metrics and reporting
Actigy BPO reports results by task and version, with sample sizes and unresolved cases. Review score distributions, not only an average. Inter-reviewer agreement measures consistent judgments; it does not prove the rubric is correct.
| Measure | Meaning |
|---|---|
| Score distribution | Counts by rubric score, task and version |
| Agreement | Matching judgments across independently reviewed examples |
| Failure categories | Errors grouped by the client's definitions, with annotated examples |
| Accepted output and rework | Reviews accepted first time and work returned for correction |
How Actigy BPO tests AI agents and chatbots
Actigy BPO runs scripted conversations and agreed exploratory tests in your test environment. Reviewers check multi-step tasks, approved tool use, escalation and handoff to people. The client sets permissions and owns integrations, fixes and deployment.
AI agent testing services
Actigy BPO records whether an agent completes the assigned task and where its actions depart from the expected path. Test cases include missing information, failed tools and requests outside the agent's authority.
Regression checks after changes
Actigy BPO repeats the agreed test set after a model, prompt or policy change. Keep versions and reference answers with each result. Compare changed behavior, investigate new failures and send evidence to the model team.
Fairness, reliability and evaluation challenges
Actigy BPO checks outputs against the client's fairness rubric and approved examples. Reviewers flag inconsistent treatment and record the evidence. Your policy owner decides how to resolve it. A passing sample does not prove that every user or situation receives the same treatment.
| Challenge | Control to agree |
|---|---|
| Changing benchmarks | Version the rubric, test set and reference answers. |
| Growing volume | Keep review depth and calibration time in the capacity plan. |
| Benchmark versus real use | Include client-approved examples of actual tasks and exceptions. |
| Subjective or disputed scores | Record reasons and route disagreement to a named owner. |
Data controls, locations and review hours
Actigy BPO limits access to approved client systems and prohibits reuse of client data for other work. Agree confidentiality, retention, return and access removal before sharing a sample.
Actigy BPO delivers from teams in Bulgaria, Romania, Poland and Ukraine. Bulgaria, Romania and Poland are European Union (EU) hubs. Ukraine is outside the EU. The written scope names work locations, subprocessors and any transfer terms.
Actigy BPO teams in Central and Eastern Europe cover the full UK business day and the US morning. Confirm live-review shifts and fallback arrangements separately from batch deadlines.
Actigy BPO uses ISO 9001-aligned quality management. Aligned, not certified. SOC 2-aligned controls: role-based access, logging and segregation of duties. Aligned, not audited. Review the security practices and data-processing terms.
LLM evaluation vendors: pricing and delivery choices
Actigy BPO prices most services per FTE by role and sends a written quote after the process audit. Compare task scope, calibration, review depth and retained client work before comparing prices.
| Responsibility | Actigy BPO managed team | In-house team | Freelancer or staffing agency |
|---|---|---|---|
| Daily management | Agreed team lead | Client manager | Confirm client supervision |
| SOP ownership | Client-owned | Client-owned | Confirm ownership in contract |
| Quality checks | Agreed QA and sampling | Client review system | Confirm review responsibility |
| Coverage | Scoped shifts | Internal schedule | Contracted availability |
| Start | Audit and paid pilot | Hiring and training readiness | Selection and onboarding |
| Price basis | Per FTE by role | Employment and operating costs | Contracted fee basis |
| Exit | Agreed records and access handback | Internal knowledge transfer | Agreed deliverables and handback |
Actigy BPO fits when you need pre-release output reviews, agent workflow tests or regression batches with approved examples and a client decision owner. Actigy BPO is not the right fit when you need model engineering, security testing or fully automated scoring without human review.
Annotation evidence, not an evaluation accuracy target
Actigy BPO's AI data annotation case study reports over 1.5 million annotations and a 99.92% QA pass rate (Actigy BPO-reported). Those figures describe light detection and ranging (LiDAR) and video annotation, not chatbot factuality or model improvement.
Results come from anonymized client engagements, are reported by Actigy BPO and are not independently audited. Results depend on scope.
LLM evaluation services FAQ
What are LLM evaluation services?
LLM evaluation services test model responses against defined quality and task criteria. Actigy BPO provides human review of outputs, response comparisons and agent task checks in approved client tools. Reviewers apply your rubric and record the reason for each score.
Your model team decides which samples represent real use. A review result describes that sample, rubric and model version; it does not establish quality across every possible request.
Why use human evaluation when automated evals exist?
Actigy BPO uses human review to examine meaning, context and policy cases that a numerical check alone does not explain. Automated metrics still help your model team compare repeatable tests. Reviewers investigate examples, compare answers and record the reason for a judgment.
Use the same task set to compare both views. Disagreement can reveal a weak rule or unclear reference answer, not just a model error.
How does Actigy BPO keep reviewer scores consistent?
Actigy BPO trains reviewers on the client's rubric and checks their scores against approved reference items. Calibration compares judgments before reviewers enter the live queue. Maker-checker quality assurance (QA) and sampling then expose drift and disputed labels.
A named client owner settles unclear rules. Track agreement by task and rubric version, with sample sizes recorded. Agreement alone does not prove that a reference answer is correct.
Can Actigy BPO test AI agents and chatbots?
Actigy BPO can run agreed agent and chatbot tasks in client-approved test environments. Reviewers follow scripted conversations and examine permitted variations, tool choices and handoffs. They record the task outcome and the step where an error occurs.
Your engineering team owns test access, integrations, fixes and release approval. Safety checks follow your policy; this service does not include penetration testing or permission to act in live accounts.
Which tools and platforms can the reviewers work in?
Actigy BPO reviewers work in the client's labeling and evaluation tools, under documented guidelines with maker-checker QA. The process audit checks access, task fields, review states and the required output format. Tool names alone do not confirm that a workflow is ready.
Your team approves each system and permission. Confirm whether reviewers can view source evidence, record comments and export accepted results before committing a batch.
How is client data protected in evaluation work?
Actigy BPO limits evaluation access to the systems and records the client approves. The written scope names work locations, roles, permitted use, retention and transfer terms. Reviewers do not reuse client data for another project.
Agree how to mask sensitive records and remove access at exit. Your team keeps raw-data hosting and release decisions. Review the proposed controls against your own data restrictions before sharing samples.
How are LLM evaluation services priced?
Actigy BPO prices evaluation work per full-time equivalent (FTE) by role after the process audit. Task length, rubric complexity, review depth and coverage determine the proposed staffing. The written quote separates agreed delivery roles from work your model team retains.
Actigy BPO uses a paid pilot to measure accepted reviews and rework for the proposed sample. Cost per accepted review is a comparison measure, not a promised per-task fee or output rate.
Related human-review workflows
Actigy BPO separates evaluation from training-data work. Review AI outsourcing and RLHF services for adjacent tasks.
Compare QA and software testing, AI plus human operations and AI automation statistics without treating research figures as service outcomes.
Page updates
October 4, 2026: published evaluation scope, calibration, agent checks, data controls and pricing boundaries.
Test one evaluation queue
Share task examples, rubric versions, review volume and the decisions your team retains.
Scope a pilot
What happens next
The team reviews the workflow before proposing a written scope. You decide whether to start a paid pilot after reviewing it.