Know where your LLM feature fails — before your customers do.

Arjunworks helps B2B SaaS teams check document-grounded AI answers before a model or prompt change. Start with 30 human-reviewed cases and a reusable evaluation set. Proposed audit fee: 15,000 SEK excluding VAT, subject to scope confirmation. Based in Uppsala, working in English.

20 minutes, no deck. If the audit doesn't fit your situation, I'll say so on the call.

Shipping was the easy part.

The feature demos well. Then it meets production:

  • Someone on the team pastes outputs into a spreadsheet before every release and eyeballs them.
  • A prompt tweak fixes the case a customer complained about — and silently breaks three others.
  • A model upgrade lands, and nobody can say whether quality went up or down.
  • The worst outputs are found by customers, not by you.

A repeatable evaluation set helps you record which answers are supported, where evidence is incomplete, and what changes between releases. It gives your team a documented basis for deciding what to check next.

Three fixed-scope services. Nothing else.

Start with a bounded audit. Any implementation work is scoped separately. Prices exclude VAT and are confirmed against the agreed scope.

LLM Reliability Audit

15,000 SEK excluding VAT · target 5 working days

A source-traceable baseline for one English-language, document-grounded AI workflow.

  • 30 agreed cases, all human-reviewed. Five are used with your reviewer to agree the scoring rubric before the remaining 25.
  • A verdict, source references or explicit evidence gaps, and uncertainty for each case.
  • A prioritised findings report, reusable case set, documented replay handover and findings call.

One fixed source snapshot and one model/prompt configuration. Document volume, formats and answer length are bounded in the scope agreement; extra cases or configurations are scoped separately.

The delivery target starts once scope, usable inputs, permitted access and commercial arrangements are agreed. Invoice and payment terms are confirmed with Frilans Finans before the final agreement.

Results describe the evaluated sample. Finding no failures is a valid result; it does not establish production-wide accuracy.

AI Reliability Pilot

from 55,000 SEK

The audit's eval set, made permanent.

  • Automated evaluation runs wired into your release process.
  • Guardrails on the failure modes the audit found.
  • Regression monitoring, so a prompt or model change that hurts quality is caught before release, not after.

Scope and price are fixed in writing after the audit, before signature.

Phase 1 Hardening

from 85,000 SEK · 15 delivery days

Implementation of the top items on your hardening roadmap.

  • Four named acceptance criteria, agreed before signature.
  • Everything not listed in the criteria is explicitly out of scope.

You get exactly what the contract names. If that sounds restrictive, it is — it's what makes the price and the deadline real.

From first call to delivered audit

  1. Intro call — 20 minutes. You describe the feature and where it hurts. I tell you whether the audit fits. If it doesn't, I'll say so and point you somewhere useful.
  2. Scope agreement. One workflow, 30 cases, source and answer-length limits, the model/prompt configuration, permitted data processing, price, any usage-cost cap, payment terms and delivery date. Agreed before work starts.
  3. Five-working-day target. Starts after scope, usable inputs, access and commercial arrangements are confirmed. We calibrate on five cases, review the remaining 25, preserve the evaluated outputs and document the findings. Waiting for agreed client inputs moves the delivery date accordingly.
  4. Walkthrough and handover. You receive the report and evaluation files, follow the replay procedure with a preserved example, and discuss the findings on a call. One consolidated round of factual or scoring corrections is included, with its feedback window agreed in the scope. Your team retains the release decision.
Book the intro call

Products by Arjunworks

Tools for inspecting AI answers and the evidence behind them.

Main product · Early-stage prototype

Evidence Audit

Examine AI-generated claims against retrieved sources, with a verdict and evidence for each claim. This is our main product development focus.

Explore Evidence Audit

Live · Access by invitation

Research Observatory

Ask questions of public research papers, inspect the retrieved passages and open the cited PDF page. A working research preview with a shared paper library.

Explore Research Observatory

Proof of work

There are no client projects to show yet — more on that below. What I can show is a working prototype of the discipline these services are built on.

Evidence Audit is an early-stage prototype that checks AI-generated claims against retrieved sources before anyone acts on them. It takes an AI answer, splits it into individual claims, checks each claim against source material, and returns a per-claim verdict — supported, partially supported, or unsupported — plus the risk if the claim is wrong.

In one illustrative run, on a question about lithium-ion battery degradation, it audited five claims from an AI-generated answer: three supported, one partially supported, one unsupported. The unsupported claim was contradicted by the retrieved sources — the kind of confident, wrong statement that gets built into downstream decisions when nobody checks. One run on one question is not a benchmark, and I won't present it as one. The code and full write-up are public:

github.com/AIArjun/evidence-audit-rd

The paid audit applies this approach through human review: check your feature's outputs against agreed sources and expected behaviour, record evidence gaps, and deliver a reusable case set. Evidence Audit remains an early prototype.

C1 — SEI layer growth accelerates at elevated temperature SUPPORTED
C2 — Electrolyte decomposition at high temperature SUPPORTED
C3 — Cathode transition-metal dissolution SUPPORTED
C4 — Loss of active material PARTIALLY SUPPORTED
C5 — Lithium plating at elevated temperature UNSUPPORTED

Output of a single illustrative run. Not a benchmark.

Why there are no client logos on this page

Because there are no clients yet. Arjunworks is new, and I'd rather say that plainly than pad this page with a logo strip.

What you get instead of references:

  • Fixed scope, in writing. Thirty cases for one workflow, with agreed source limits and one model/prompt configuration.
  • Clear service fee. 15,000 SEK excluding VAT, subject to confirming the bounded scope.
  • Defined acceptance. All 30 cases are recorded, findings are traceable to evidence or explicit gaps, the report and call are delivered, and your reviewer can follow a preserved example.
  • Controlled usage costs. Any additional model or hosting spend has a named payer and an agreed numeric cap before paid runs, including runs through your own accounts.

References will appear here when they exist and their owners agree to be named. Not before.

Who you'd be working with

Arjunworks is Arjun Ponnaganti, working from Uppsala in English.

  • MSc in AI & Machine Learning, Uppsala University.
  • Co-author of four peer-reviewed publications in applied machine learning (2024, under a prior affiliation with VISTAS, Chennai), spanning computer vision, network optimisation, medical ML and regression modelling. None of them is about LLM evaluation. What they demonstrate is applied ML work done rigorously enough to survive peer review.

    The two lead publications:

    • Ponnaganti Arjun et al. (first author), "Enhancing House Price Prediction Accuracy and Precision: A Data Mining Approach with Python and Stacking Algorithm", Finance and Law in the Metaverse World, Springer, 2024. DOI 10.1007/978-3-031-67547-8_17
    • "An Approach to Radiotherapy Treatment Planning based on Machine Learning Algorithms", 2024 IEEE International Conference on Expert Clouds and Applications (ICOECA). DOI 10.1109/ICOECA62351.2024.00143
    • Plus two further publications, IJISAE, 2024.
  • AI systems built with Python, FastAPI, RAG, agents, automation.
  • Selected among 60 participants from 300+ European applicants, SYE Hackathon 2026, competing as Arjunworks.

How your data is handled

Before you transfer documents, logs or examples, we agree the processing environment, permitted access, any providers involved, and retention and deletion arrangements. Your confidential material does not enter the shared Research Observatory demonstration. Evaluation can use preserved outputs or your approved environment; any new model runs require agreed access and a usage cap. Production write access is unnecessary.

Questions worth asking

What do you need from us for the audit?

Approved documents, sample questions and expected behaviour, actual outputs, and retrieved passages or logs where available. We agree document volume, formats and answer length before work starts. One technical reviewer helps calibrate the rubric on five of the 30 cases and reviews the handover. Sources and the processing environment are agreed before data transfer.

What if we don't have 30 suitable cases?

Bring the examples you have to the scoping call. We assess whether a useful 30-case set can be agreed within the scope. If the available material needs more preparation or a different scope, we agree that before quoting or starting work.

How do we agree that the audit is delivered?

Every case has a preserved output, verdict, evidence references or a documented gap, and uncertainty. Your reviewer can follow the handover to review at least one preserved case, and you receive the report and findings call. One consolidated correction round is included, with a feedback window agreed in the scope.

Acceptance does not depend on finding a particular number of failures. The sample is not a production-wide accuracy benchmark, and a new model run may produce a different answer.

What is outside the audit scope?

Application building, remediation, deployment, new evaluation-platform or CI integration, model training, continuous monitoring, certification and specialist medical, legal or safety judgments. Additional cases, configurations or implementation work require a separate scope agreement.

How do payment and invoicing work?

The proposed audit service fee is 15,000 SEK excluding VAT, subject to scope confirmation. Invoicing is proposed through Frilans Finans; the order form and permitted payment terms must be confirmed before the final agreement. No deposit schedule or prepaid discount is currently offered.

Any additional model or hosting costs need an agreed numeric cap and named payer before paid runs, including runs through customer accounts.

Will you sign an NDA?

Yes.

Do you work in Swedish?

The working language is English.

What happens after the audit?

One of three things: the feature is in decent shape and you keep the eval set; you take the roadmap and implement it in-house; or we scope a pilot. All three are fine outcomes.

What stack do you work with?

The audit is stack-agnostic — if the feature's inputs and outputs are visible, it can be evaluated. Pilots integrate with whatever CI you already run.

We're not a SaaS company. Can we still book?

The services are built for B2B SaaS teams shipping LLM features, because that's where this format fits best. If you're outside that and still want a call, book one — you'll know within 20 minutes whether it makes sense, and if it doesn't, I'll say so.

Book the intro call

20 minutes. You talk about the feature; I tell you whether the audit fits. If it doesn't, you leave with a pointer, not a pitch.

or email ponnagantiarjun644@gmail.com