OUR PRODUCT

AI-Powered Sentencing Analytics

Our platform helps public defenders, prosecutors, judges, researchers, and advocates, regardless of technical background, investigate how California sentences people. It draws on a database of more than 90,000 individual prison sentences, sourced from the California Department of Corrections and Rehabilitation (CDCR) through public records requests.

Describe who you are looking for in ordinary language, by controlling offense, prior convictions, sentence length, time served, or sentencing county, and the platform builds a matching cohort. From there you can compare outcomes across racial groups, measure disparities, and export the results as statistical evidence.

Statement on AI Use:

Statement on Privacy:

How It Works

Process:

  1. You ask. Pose a question the way you would describe a case to a colleague. Example: “Show me everyone sentenced for a robbery-related offense.”
  2. The platform interprets. A large language model identifies the relevant field (Offense Description), extracts the search terms (theft, robbery), and sets the logic that connects them (OR). It shows you what it understood before it runs, so nothing is a black box.
  3. It builds the cohort. Your criteria run against all sentencing records and return every matching sentence.
  4. It measures the disparity. The model breaks the cohort down by race and other attributes and quantifies the differences in sentencing outcomes using statistical techniques like odds ratios.
  5. You export. Carry the visualizations, cohort lists, and statistics into a motion, a petition, or a research paper.

Use Cases

Build statistical evidence for the racial justice act:

A single query reveals overrepresentation of Black individuals in robbery and theft offenses relative to every other group
Chart showing Odds Ratio formula with example values comparing Black and White groups and interpretation guide for OR results.
Built-in odds ratio calculations quantify disparity in third-strike sentencing for Black individuals, statistical evidence that can support an RJA prima facie showing

The California Racial Justice Act prohibits a sentence imposed on the basis of race, and it lets a defendant prove a violation with statistical evidence: a showing that people of their race received longer or more severe sentences than similarly situated people of other races convicted of the same offense.

We help produce analyses for prima facie showings and discovery motions (prior to an evidentiary hearing). Point the tool at an offense and it breaks the cohort down by race and surfaces the disparity. We use statistical methods such as odds ratios, relative risk, and chi-square tests, measures recognized in prior RJA cases like People v. Windom.

Because the RJA makes statistical and aggregate data admissible, and does not require statistical significance to establish a difference, an analysis like this can support the prima facie showing that triggers an evidentiary hearing. Work that once required a statistician and weeks of effort, an attorney can now generate in minutes.

find similarly situated cases:

Select and categorize penal codes to search for similarly situated cases in the population
Build a cohort of cases to use for racial disparity analysis

Every disparity claim under the Racial Justice Act rests on a comparison between similarly situated people: individuals convicted of the same offense with comparable records. Assembling that comparison group by hand is slow and error-prone.

We automate this cohort building task. Sort penal codes into predefined offense types or define your own, and it groups together the individuals who share the same profile of current and prior commitments. The result is a defensible comparison cohort, the “similarly situated” population the statute requires, built in seconds rather than weeks of manual case review. The same capability surfaces candidates for second-look and resentencing review.

Tutorials & Resources

Talks, guides, and the research grounding our platform design.

 

Open Code, Open Data

The full platform is not open source, but the models, algorithms, and statistical methods behind it are public on GitHub, along with evaluations of our AI integrations benchmarked against human statisticians.

The statistical engine behind our disparity analyses the same measures accepted in prior RJA cases

Vector-based methods for identifying similarly situated cases, the comparison group the RJA requires

How we test our AI: evaluations of the platform’s LLM integrations, benchmarked against analyses produced by human statisticians.
The modular Python framework that builds “similarly situated” comparison cohorts from sentencing records.

Tutorials

If you make use of our dataset(s) or code, please cite our work as follows:

APA style:
BibTex:

Citation

If you make use of our dataset(s) or code, please cite our work as follows:
APA style:
BibTex:

License

All content in this project is licensed under GNU AGPLv3  License: AGPL v3LicenseAGPL v3
You must give appropriate credit to our work. You may not use our work for commercial purposes, which means anything primarily intended for or directed toward commercial advantage or monetary compensation.