JevSDSQL:

A self-developing SQL database with JEV based semantic operators and natural language queries

41 JEV based semantic operators and conditional workflows cover classification, extraction, ranking, matching and verification. Parallel evaluation and reusable evidence reduce repeated work. English and Simplified Chinese queries, optional LLM planning, and text-to-table extraction make the database accessible through a web workspace and HTTP API.

The operators combine JEV's native Noul, Choice and Score primitives with database and workflow logic. JEV is provided by TypeSafe; PostgreSQL handles storage, joins and arithmetic.

Installation · User guide · 简体中文 · Operator reference

Features

  • Query data in English or Simplified Chinese. Inspect the SQL, correct its interpretation and rerun it from query history.
  • Choose JEV planning or hybrid planning. In hybrid mode, JEV selects relevant context, an LLM proposes SQL, and JEV reviews the proposal.
  • Filter and classify text by meaning. Save reviewed definitions as reusable features, such as whether a message requests action.
  • Extract database entries from documents. Describe the rows and columns, then review typed values alongside their source text before importing.
  • Preview inserts, updates and deletes before committing them.
  • Call semantic operators for extraction, ranking, matching, verification and conditional workflows.

The self-developing part is the semantic layer: definitions, evidence and corrections can be saved, reviewed, reused and refreshed as data changes. New concepts require approval before promotion.

Getting started

You need Python 3.11 or newer, PostgreSQL and a TypeSafe API key. Python 3.13 and PostgreSQL 17 are the verified configuration. Hybrid mode also requires a configured LLM provider.

git clone https://github.com/Sheltercosmo/JevSDSQL.git
cd JevSDSQL
python -m venv .venv

Activate the environment with .venv\Scripts\Activate.ps1 in PowerShell or source .venv/bin/activate in Bash, then install:

python -m pip install -r requirements.lock.txt
python -m pip install --no-deps -e .

Copy .env.example to .env and follow the database setup to create the application role and configure credentials. Then run:

python -m scripts.migrate_generic
python -m sdd.cli serve

Open the English workspace or Simplified Chinese workspace. Connect with your database access token and import data through Manage data.

Query your data

Select a dataset and enter a question such as:

For each supplier, show the total quantity delivered, largest total first.

Choose Preview plan to inspect the SQL. You can edit a planning decision, select an alternative interpretation or edit the SQL directly. Run the query when the proposal matches your intent. Data changes require a separate commit.

The same workflow is available through POST /ask:

{
  "question": "For each supplier, show the total quantity delivered, largest total first.",
  "dataset_ids": ["deliveries"],
  "planner_mode": "hybrid",
  "execute": false
}

Use an ID or name from your catalog in dataset_ids. Set planner_mode to jev to plan without LLM generation. See the query API guide and the local API reference for authentication and complete requests.

Semantic operators

The API provides 41 operators and WORKFLOW. For example, send this request to POST /jev/call to test a proposition:

{
  "operator": "JEV.NOUL",
  "arguments": {
    "state": "The shipment arrived on Tuesday.",
    "proposition": "The shipment has arrived."
  },
  "limits": {"max_judgments": 1, "max_requests": 1}
}

Results distinguish a known value, an unknown answer and work that was not evaluated. Independent judgments can run in parallel; compatible evidence can be reused.

Use the operator guide to select an operator and set budgets. The function reference includes arguments, examples and result contracts for every operator.

Performance and limitations

For this release, a local comparison on 100 BIRD Challenging questions produced the following results. One reference timed out, leaving 99 scored questions. Held proposals were included.

Method Matching SQL answers Median request time Estimated cost per 100 attempts
JEV 20/99 8.55 s $0.389
LLM baseline 39/99 8.68 s $3.262
Hybrid 34/99 14.23 s $2.839

Timing covers 94 questions run with up to three concurrent cases. Costs use recorded usage and fixed assumed rates; they are not current prices or subscription charges. See performance and cost for the measurement conditions and missing usage.

Hybrid cost less on this sample but did not outperform the LLM baseline in accuracy or speed. JEV planning supports a bounded set of query structures. All modes can produce incorrect interpretations; inspect proposals and result completeness before relying on them. JEV and LLM model weights are not included.

Documentation and contributions

See the documentation index for text imports, semantic features and architecture. To report a problem, include a small synthetic dataset, the request and the expected result. Development setup and checks are in CONTRIBUTING.md.