Comprehensive Skills Assessment Platform

Every answer tested. Every point earned.

cSAP runs programming, SQL and open-ended exams in the browser and grades them for you: code in many languages against test cases generated for every submission, written answers by an AI grader against your rubric, with the evidence behind every point. It connects to your student information system, and AI agents can run it over MCP.

One private deployment per institution

  • 1answer key for every programming language
  • 3exam types: code, SQL and written
  • 3interface languages: EN, AZ, RU
  • 0installs: students need only a browser

Three exam types

Code, queries and written answers. One platform grades them all.

Each exam is one type. Pick the one that tests the skill you teach, and cSAP checks every answer the same way for every student.

Programming

Students read input and print output in the language they choose at the start. You write one reference solution and one test-case generator, and they grade every language the exam offers.

  • Python
  • C++
  • Java
  • C#
  • JavaScript
  • Go
  • Rust
  • Test cases generated for every submission
  • A time limit on every case
  • A wrong answer shows the input and the first line that differs

SQL

Multi-step tasks where every statement counts. The reference script and the student's run in separate in-memory SQLite databases, from the same setup script.

  • Query results compared row by row
  • Other statements judged by the database they leave behind
  • Row order matters only when the task asks for it

Open-ended

Short written answers, scored by an AI grader against your rubric. A point counts only when the grader backs it with a quote that is really in the answer, and teachers can override any point.

  • Rubric points, added up in code
  • Word, character and answer-language limits
  • One submission per question, replaceable while it waits

How it works

From a spreadsheet to graded results

  1. 01

    Write the exam in a spreadsheet

    Each row is a problem: its description, a reference solution, a test-case generator and a time limit. A written question gets a reference answer and a rubric instead. Start from the templates and attach images as assets.

  2. 02

    Students sit it in the browser

    They sign in with the assessment name and their personal code. No accounts, nothing to install. Each student draws their own set of problems, picks a language when the exam offers several, and works in a full code editor that starts from a template for that language.

  3. 03

    Grading runs in the background

    Every submission goes into a queue for an isolated grader. The page updates by itself, so students can move on to the next problem while they wait. Written answers also show their place in the queue.

  4. 04

    Review and share

    Results come together in one table, with attendance, the average score and each student's focus-loss count. Open any student's submission history, share a read-only link or print the results.

Integrity

Built for exams that count

No platform can promise that nobody will cheat. cSAP makes the shortcuts hard to take and easy to see, in the grader and in the exam room.

How a programming answer is graded

  1. 1

    Cases generated every run

    The generator runs again for every submission. When it draws random input, as the templates do, memorized outputs do not carry over.

  2. 2

    Reference first, then gone

    The reference solution runs first. Its output stays in the grader's memory, and its files are deleted before the student's code is written.

  3. 3

    Student code runs alone

    Once per test case, as an unprivileged user, under a time limit.

  4. 4

    The grader decides

    Pass or fail is a verdict the grader sets after comparing outputs. There is no success message a program could print to fake it.

And for written answers

Every awarded point needs a quote that really appears in the answer, the total is added up in code rather than by the model, and results are signed so a forged grade is dropped.

Exam-room controls

  • Focus tracking

    Leaving the exam window is counted. At the limit, the student is blocked.

  • Copy, paste and undo off

    Two switches per assessment: one blocks copy, cut and paste, the other undo and redo.

  • One device

    Students stay on the device they started on. Teachers can reset it.

  • Kiosk mode

    A ready-made command opens cSAP in Chrome's full-screen kiosk mode on Windows.

  • Personal problem draw

    Every student draws their own problems from your pool, by score.

  • Time window

    Set a start and an end. Students see the countdown.

  • Hidden titles

    Show problems by number only.

  • Limited resets

    Decide how many times a student may draw a new problem list.

AI grading

Written answers, graded with evidence

Open-ended questions are scored by a language model against the rubric you write. cSAP checks every quote and adds up the score, and a teacher can change any point.

  • Your rubric, point by point

    One line per point: a score and the statement it rewards. The points add up to the question's score.

  • No quote, no point

    For every point it awards, the grader must quote the answer. A quote that is not in the answer is thrown out, and the point with it.

  • Arithmetic in code

    The model judges each point; cSAP adds them up. The total never comes from the model.

  • Teachers have the last word

    Override any point, or regrade one answer or a whole question after editing the rubric. Overrides record who made them and when.

  • Tested against tricks

    Prompt injection, keyword stuffing and swapped claims are part of the test set every supported model is checked against.

  • Answer keys stay private

    The reference answer and the rubric never reach students, and the per-point feedback is for teachers only.

The supported grader today is an open-weight model, run through a hosted provider with data collection turned off. Running a model on your own server with Ollama is built in, for models that pass the same test set.

An answer graded against a three-point rubric. Point at a rubric line to find its evidence.

Q2Why can deep recursion cause a stack overflow?

Student answer

Each recursive call adds a new frame to the call stack, and the frames are only removed when the calls return. If the recursion goes too deep, or there is no base case, so it never stops, the stack runs out of space.

Rubric

  1. Each call adds a stack frame

    “adds a new frame to the call stack”

    2
  2. A missing or unreachable base case

    “no base case, so it never stops”

    1
  3. Names a fix: iteration or a smaller depth

    No evidence in the answer

    1
Score 3 / 4

Multilingual

In the language your students think in

Students switch the interface and the problem text between English, Azerbaijani and Russian. Translations are optional and fall back to the original field by field, so a half-translated exam still works.

  • Code and test cases are written once. Only the prose is translated, so grading never changes.
  • Written exams can require answers in a chosen language.

The same problem in three languages

English

Pair with the given sum

Read n numbers and a target. Print the positions of two numbers that add up to the target.

Input
5
3 9 1 7 4
8
Output
2 3

Azərbaycanca

Cəmi verilən ədədə bərabər olan cüt

n ədəd və bir hədəf oxuyun. Cəmi hədəfə bərabər olan iki ədədin mövqelərini çap edin.

Giriş
5
3 9 1 7 4
8
Çıxış
2 3

Русский

Пара с заданной суммой

Считайте n чисел и целевое значение. Выведите позиции двух чисел, сумма которых равна целевому значению.

Ввод
5
3 9 1 7 4
8
Вывод
2 3

Integrations and AI agents

Connected to your systems and your AI agents

Final scores flow back to the system that keeps your grades, and AI agents can run cSAP for you over MCP. Both go through the same API, with the same role checks as the admin pages.

Sync scores with your student information system

Import students and subjects each semester. Your system then collects each student's final score through the integration API and marks the assessment as synced, so nobody types scores in by hand.

  • Students and subjects imported per semester from a spreadsheet
  • Each student's final score, with the maximum score and the date
  • Assessments filtered by synced or not yet synced

Built around one SIS so far, so connecting yours may take some adapting.

Your AI agents, over MCP

cSAP is also a Model Context Protocol (MCP) server. Connect any agent or AI platform that reaches MCP servers over HTTP with a token, then ask in plain language: it creates assessments from a workbook, reads results and manages users and integration files.

  • Every tool is an API endpoint, with the same validation and role checks
  • Give each agent its own account, with only the roles it needs
  • Choose which parts of the API become tools
  • REST API

    Assessments, users and results over HTTP, documented with OpenAPI.

  • One token for both

    The REST API and the MCP server accept the same token, in an Authorization or x-api-key header.

  • Files in the tool call

    Agents upload exam workbooks and images inside the tool call, the same way the admin pages do.

For teachers and admins

Everything around the exam, too

  • Results table

    Totals, attendance, the average score and each student's passed and failed counts. Share a read-only link or print it.

  • Submission history

    Each student's attempts, with the code or the answer and its verdict.

  • Exams as workbooks

    Download a template, add one row per problem and upload it with its images.

  • Roles

    Separate rights for user managers, integrations and submission analysts.

  • Submission log

    Every passing submission is written to a log that submission analysts can read, or download as one archive.

  • Your own deployment

    Each institution runs its own copy, with its own data and graders. Docker Compose, nothing exotic.

FAQ

Questions, answered

What kinds of exams can I run?

Three kinds. Programming exams, answered in any of the languages you offer. Multi-step SQL exams, checked against a database. Open-ended exams with short written answers, scored against your rubric. Each exam is one kind.

Which programming languages are supported?

Python, C++, Java, C#, JavaScript, Go and Rust. You choose which of them an exam offers. Each student picks one before starting, and can change it only by resetting their problem list, which you limit.

How is a programming answer checked?

You provide a reference solution and a test-case generator. For every submission, cSAP runs the generator, runs the reference to get the expected output, then runs the student's program on the same input and compares the two. Every case has a time limit, and a generator that draws random input gives every submission new cases.

Can students hard-code the expected output?

In programming exams it is hard to make that pay off. When the generator draws random input, every submission meets new test cases; the reference solution's files are gone before student code runs, and that code runs as an unprivileged user. In the browser, focus tracking, the copy-paste and undo switches and the one-device check make it harder still.

Can I trust AI grading?

You stay in charge of it. The grader gives a verdict on each rubric point and must quote the answer for every point it awards. Quotes that are not in the answer are rejected, and the score is added up by cSAP, not by the model. Teachers see the evidence for each point and can override it or regrade.

Do students need an account?

No. They sign in with the assessment name and the personal code you give them. A browser is all they need.

Does cSAP work with our student information system?

It can. cSAP imports students and subjects each semester, and your system collects final scores through an integration API. That API was built around one system, so connecting yours may take some adapting. The REST API and the MCP server cover anything else you want to automate.

Can an AI agent work with cSAP?

Yes. cSAP serves its admin API as an MCP server over HTTP. Give the agent its own account, with MCP access and only the roles it needs, and it can create assessments, upload workbooks, read results and manage users. Every tool call goes through the same checks as a REST API call.

Where does cSAP run, and who sees our data?

Each institution gets its own deployment with its own database, and programming and SQL answers are graded inside it. Written answers go to an open-weight model at a hosted provider, with data collection turned off. Running the model on your own server with Ollama is built in, for models that pass our test set.

How do I create an exam?

Fill in an Excel template: one row per problem or question, and a Students sheet that sets how many problems of each score every student draws. Upload it with its images, and set the time window and the exam-room controls.

Bring us one of your exams

Send a problem or a question you set this term. We will set it up in cSAP and show you how it is graded: test cases, rubric and all.

or write to sales@csap.app