Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
45 changes: 45 additions & 0 deletions .changeset/20889-analytics-native-measure-number.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
---
'@objectstack/core': minor
'@objectstack/driver-sql': patch
'@objectstack/service-analytics': patch
---

fix: the analytics native-SQL path answers a measure its response declares `number` as a number on every dialect, presented by the one rule `driver-sql`'s `aggregate()` applies, which `@objectstack/core` now exports as `AGGREGATE_ANSWER_KIND` and `presentAsNumber` (#20889)

Clause-②: yes (widening)

**New exports.** `@objectstack/core` exports two names, moved here unchanged
from `@objectstack/driver-sql`, which now imports them instead of keeping them
private:

- `AGGREGATE_ANSWER_KIND`: what each declared aggregate function answers.
`count`, `count_distinct`, `sum` and `avg` answer `'number'`; `min` and `max`
answer `'column'`, a value of the aggregated column.
- `presentAsNumber(value)`: the `'number'` presentation. A string `Number()`
reads as a number becomes that number. Any other value is returned as given:
a number, `null`, a boolean, empty or blank text, or text that reads as NaN.

**What changed.** On PostgreSQL, `POST /api/v1/analytics/query` and
`POST /api/v1/analytics/dataset/query` answered through `NativeSQLStrategy`
returned count, count_distinct, sum, avg, and min / max over a numeric column
as strings, such as `count: "2"` and
`sum: "500.000000000000000000000000000000"`, while `fields[]` declared
`number`. A dataset's `row_count` did the same, and a measure-scoped count
mixed `"1"` with the number `0` in one column. SQLite answered numbers. The
strategy now presents each measure column by its declared aggregate function,
through the same table and presenter as `SqlDriver.aggregate()`. `min` / `max`
are presented only when their column is declared numeric, so `max` over a text
column, every dimension, and expression measures keep the value the database
returned.

**Precision.** The answer is one JS number, the policy `driver-sql`'s
`aggregate()` already applies. A total that needs more digits than a double
holds, such as `9007199254740993`, answers the nearest double
(`9007199254740992`), which is also what SQLite and the engine path answer.

**What did not move.** `@objectstack/driver-sql`'s behaviour is unchanged: its
`aggregate()` reads the same table, and its read presenter calls the same
function. The answers on SQLite are byte-identical. The arithmetic of the
analytics native statement did not change either. On PostgreSQL its `sum` and
`avg` still add exact decimals, so `0.1 + 0.2` answers `0.3` where the engine
path answers `0.30000000000000004`.
3 changes: 3 additions & 0 deletions packages/core/src/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,9 @@ export * from './utils/datetime.js';
// runtime dependencies, and a copy per face is how one `sum` came to answer
// two doubles.
export * from './utils/compensated-sum.js';
// [#20889] What an aggregate ANSWERS and its `'number'` presenter, moved from
// `driver-sql` so the analytics native-SQL face presents with the same rule.
export * from './utils/aggregate-answer.js';

// Export the shared batched-write helper (framework#2678)
export * from './utils/bulk-write.js';
Expand Down
76 changes: 76 additions & 0 deletions packages/core/src/utils/aggregate-answer.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license.

/**
* [#20889] `AGGREGATE_ANSWER_KIND` and `presentAsNumber` — what an aggregate
* ANSWERS, and the one `'number'` presenter every face that hands a SQL
* client's aggregate row to a caller applies: `driver-sql`'s own `aggregate()`
* (#20335) and `service-analytics`' native-SQL face (`NativeSQLStrategy`).
*
* The wire strings below are the values measured on live PostgreSQL 16.13
* through the analytics native-SQL path (node-postgres parses `bigint` and
* `numeric` to strings), and the large totals are the ones #20335 pinned at the
* driver door. Each face pins the same presentation through its own door, in
* its own package, on SQLite and PostgreSQL.
*/

import { describe, it, expect } from 'vitest';
import { AggregationFunction } from '@objectstack/spec/data';
import { AGGREGATE_ANSWER_KIND, presentAsNumber } from './aggregate-answer';

describe('[#20889] AGGREGATE_ANSWER_KIND — what each declared aggregate function answers', () => {
it('has exactly one row per declared aggregate function', () => {
expect(Object.keys(AGGREGATE_ANSWER_KIND).sort()).toEqual([...AggregationFunction.options].sort());
});

it('count, count_distinct, sum and avg answer a number; min and max answer a value of the column', () => {
expect(AGGREGATE_ANSWER_KIND).toEqual({
count: 'number',
count_distinct: 'number',
sum: 'number',
avg: 'number',
min: 'column',
max: 'column',
});
});
});

describe("[#20889] presentAsNumber — the 'number' presenter", () => {
it('turns the numeric text a SQL client hands back into the number it spells', () => {
// count / count_distinct (`bigint`), sum over an integer column (`bigint`).
expect(presentAsNumber('2')).toBe(2);
expect(presentAsNumber('11')).toBe(11);
// avg over an integer column (`numeric`, 16 places).
expect(presentAsNumber('3.5000000000000000')).toBe(3.5);
// sum / avg / min / max over the exact-decimal column (`numeric(65,30)`).
expect(presentAsNumber('500.000000000000000000000000000000')).toBe(500);
expect(presentAsNumber('20.500000000000000000000000000000')).toBe(20.5);
expect(presentAsNumber('0.125000000000000000000000000000')).toBe(0.125);
expect(presentAsNumber('-7.250000000000000000000000000000')).toBe(-7.25);
});

it('precision policy: a total a double cannot hold answers the nearest double, never a string', () => {
const big = presentAsNumber('9007199254740993.000000000000000000000000000000');
expect(typeof big).toBe('number');
expect(big).toBe(9007199254740992);
expect(big).toBe(Number('9007199254740993'));
const decimal = presentAsNumber('12345678901234567.123456789000000000000000000000');
expect(typeof decimal).toBe('number');
expect(decimal).toBe(12345678901234568);
expect(decimal).toBe(Number('12345678901234567.123456789'));
});

it('passes every value that is not numeric text through as given', () => {
expect(presentAsNumber(7)).toBe(7);
expect(presentAsNumber(0.30000000000000004)).toBe(0.30000000000000004);
expect(presentAsNumber(null)).toBeNull();
expect(presentAsNumber(undefined)).toBeUndefined();
expect(presentAsNumber(true)).toBe(true);
expect(presentAsNumber('')).toBe('');
expect(presentAsNumber(' ')).toBe(' ');
// PostgreSQL's `numeric` 'NaN', and text `Number()` cannot read, stay as written.
expect(presentAsNumber('NaN')).toBe('NaN');
expect(presentAsNumber('abc')).toBe('abc');
const date = new Date(0);
expect(presentAsNumber(date)).toBe(date);
});
});
120 changes: 120 additions & 0 deletions packages/core/src/utils/aggregate-answer.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,120 @@
// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license.

/**
* [#20889] What an aggregate ANSWERS, and the `'number'` presenter that gives a
* count or a total its type — defined once, for every face that hands a SQL
* client's aggregate row to a caller.
*
* ## Why it lives here
*
* Two faces run an aggregate statement on a SQL client and must present its
* answer the same way:
*
* - `@objectstack/driver-sql`'s native `aggregate()` (#20335), where this rule
* was written: `aggregate()` reads {@link AGGREGATE_ANSWER_KIND} to choose
* the result columns that take the `'number'` presentation, and
* `presentReadValue`'s `'number'` arm is {@link presentAsNumber};
* - `@objectstack/service-analytics`' native-SQL face
* (`NativeSQLStrategy.execute`), which runs its own statement through the
* host's raw-SQL bridge — `engine.execute`, or an embedder's own
* `executeRawSql` — and so never passes through the driver's `aggregate()`.
* On PostgreSQL it answered `count: "2"` under a `fields[]` that declared
* `number`.
*
* `service-analytics` holds `driver-sql` as a development dependency only — it
* is multi-driver, and its native path also runs over an embedder's raw SQL —
* and this is the package both already stand on, the reason `compensatedSum`
* lives here too. A second transcription is how a face comes to answer its own
* type again.
*
* The table and its docblock below moved here from `driver-sql` unchanged, so
* they speak of `SqlDriver.aggregate` and `formatOutput` in that driver's
* voice.
*/

import type { AggregationFunction } from '@objectstack/spec/data';

/**
* [#20335] What each declared aggregate function ANSWERS — a derived `number`,
* or a value OF the aggregated column — and therefore which read presentation
* {@link SqlDriver.aggregate} gives its result column.
*
* - `'number'` — `count`, `count_distinct`, `sum`, `avg`. A count or a total is
* a number whatever the column held, and it is presented as one (`'number'`,
* the presenter `formatOutput` applies to a numeric field on a `find()` row).
* - `'column'` — `min`, `max`. The answer is one of the column's own values, so
* it takes that column's presentation ({@link SqlDriver.readPresentationKind}),
* exactly as before this table existed.
*
* Why the `'number'` half needs presenting at all: the SQL client hands a
* result back as the wire type of the SQL expression, not as the platform's
* value type. Measured on live PostgreSQL 16.13 and MySQL 8.0.46 through this
* driver's own connections: node-postgres parses `bigint` (OID 20 — `count`,
* and `sum` over an integer column) and `numeric` (OID 1700 — `sum` / `avg`
* over the exact-decimal numeric family, `avg` over an integer column) to
* STRINGS (`"2"`, `"500.000000000000000000000000000000"`), and mysql2 does the
* same for `DECIMAL` (`SUM` / `AVG`; its `COUNT` arrives as a number). The
* engine's rows path (`objectql`'s `in-memory-aggregation.ts`) and SQLite answer
* numbers for the same query, so `having { n: { $in: [2] } }` kept c1, c2 on
* those and no group on PostgreSQL's native path.
*
* Keyed on the function the query ASKED for, never on whether a value looks
* numeric, and deliberately not gated by dialect: the presenter only rewrites a
* STRING, so a client that already answers a number (better-sqlite3, mysql2's
* `COUNT`) passes through untouched — measured byte-identical on SQLite — and
* no list of "string-answering dialects" exists to drift.
*
* ## The precision policy — one JS number, the loss declared
*
* The answer is `Number(text)`: an IEEE-754 double, on every dialect. A `sum` /
* `avg` over the exact-decimal column (`numeric(65,30)` / `DECIMAL(65,30)`)
* whose value needs more than a double's ~15-17 significant digits, or an
* integer at or above 2^53, is ROUNDED to the nearest double — declared, not
* silent: it is the same bound `formatOutput` already puts on a `find()` read of
* that column (#16318, `valueSchemaFor`'s `z.number().finite()`, ADR-0104 D1),
* and the bound the rows path has always had (`toNumber` sums JS doubles). A
* value-dependent type — a number when it fits, a string when it does not — was
* rejected: it would reopen this defect for exactly the large totals, where a
* `having` `$in` or a chart silently stops matching. Only a string `Number()`
* reads as NaN (PostgreSQL's `numeric` `'NaN'`) is left as written, the
* presenter's existing rule.
*
* A `Record` over `AggregationFunction` on purpose: a function that joins the
* declared vocabulary without an answer here fails `tsc` rather than reaching a
* caller unpresented.
*/
export const AGGREGATE_ANSWER_KIND: Readonly<Record<AggregationFunction, 'number' | 'column'>> = {
count: 'number',
count_distinct: 'number',
sum: 'number',
avg: 'number',
min: 'column',
max: 'column',
};

/**
* [#20335, #20889] The `'number'` presentation: a STRING `Number()` reads as a
* number becomes that number; every other value — a number, `null`, a boolean,
* empty or blank text, text `Number()` reads as NaN (PostgreSQL's `numeric`
* `'NaN'`) — is returned as given.
*
* It is the body of `driver-sql`'s `presentReadValue` `'number'` arm, moved
* here unchanged: the presenter `formatOutput` applies to a declared numeric
* column on a `find()` row (#16318), and the one {@link AGGREGATE_ANSWER_KIND}
* names for a count or a total. The answer is one JS double — the precision
* policy that table's docblock states.
*
* ⛔ Ask it by a column's DECLARED meaning — the aggregate function the query
* asked for, or a column declared numeric — never because a value looks
* numeric: a text column holding `'007'` is text.
*/
export function presentAsNumber(value: unknown): unknown {
// Only strings are repaired, exactly as in `formatOutput`: a fresh
// REAL/INTEGER column already yields a number, and genuinely
// non-numeric legacy junk is left intact rather than turned into NaN.
if (typeof value === 'string' && value.trim() !== '') {
const n = Number(value);
if (!Number.isNaN(n)) return n;
}
return value;
}
81 changes: 14 additions & 67 deletions packages/drivers/driver-sql/src/sql-driver.ts
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,10 @@ import { parseAutonumberFormat, renderAutonumber, resolveAutonumberFormat, readA
// "the protocol has no such function" refusal cannot drift from what
// `AggregationNodeSchema.function` actually admits.
import { AggregationFunction, emptyGroupValueFor } from '@objectstack/spec/data';
// [#20335, #20889] What each aggregate function ANSWERS, and the `'number'`
// presenter its counts and totals take — defined once in core, for this
// driver's `aggregate()` and the analytics native-SQL face alike.
import { AGGREGATE_ANSWER_KIND, presentAsNumber } from '@objectstack/core';
import { STRUCTURED_JSON_TYPES, FILE_REFERENCE_TYPES, MULTI_OPTION_TYPES, NUMERIC_VALUE_TYPES, isMultiValueField } from '@objectstack/spec/data';
// [#16318] The per-field-type physical representation of the NUMERIC family.
// `os generate migration` reads the SAME table, in both of its formats — that
Expand Down Expand Up @@ -1530,63 +1534,11 @@ const SQL_AGGREGATE_FUNCTIONS: ReadonlyMap<string, SqlAggregateLowering> = new M
['count_distinct', { sql: 'count', distinct: true }],
]);

/**
* [#20335] What each declared aggregate function ANSWERS — a derived `number`,
* or a value OF the aggregated column — and therefore which read presentation
* {@link SqlDriver.aggregate} gives its result column.
*
* - `'number'` — `count`, `count_distinct`, `sum`, `avg`. A count or a total is
* a number whatever the column held, and it is presented as one (`'number'`,
* the presenter `formatOutput` applies to a numeric field on a `find()` row).
* - `'column'` — `min`, `max`. The answer is one of the column's own values, so
* it takes that column's presentation ({@link SqlDriver.readPresentationKind}),
* exactly as before this table existed.
*
* Why the `'number'` half needs presenting at all: the SQL client hands a
* result back as the wire type of the SQL expression, not as the platform's
* value type. Measured on live PostgreSQL 16.13 and MySQL 8.0.46 through this
* driver's own connections: node-postgres parses `bigint` (OID 20 — `count`,
* and `sum` over an integer column) and `numeric` (OID 1700 — `sum` / `avg`
* over the exact-decimal numeric family, `avg` over an integer column) to
* STRINGS (`"2"`, `"500.000000000000000000000000000000"`), and mysql2 does the
* same for `DECIMAL` (`SUM` / `AVG`; its `COUNT` arrives as a number). The
* engine's rows path (`objectql`'s `in-memory-aggregation.ts`) and SQLite answer
* numbers for the same query, so `having { n: { $in: [2] } }` kept c1, c2 on
* those and no group on PostgreSQL's native path.
*
* Keyed on the function the query ASKED for, never on whether a value looks
* numeric, and deliberately not gated by dialect: the presenter only rewrites a
* STRING, so a client that already answers a number (better-sqlite3, mysql2's
* `COUNT`) passes through untouched — measured byte-identical on SQLite — and
* no list of "string-answering dialects" exists to drift.
*
* ## The precision policy — one JS number, the loss declared
*
* The answer is `Number(text)`: an IEEE-754 double, on every dialect. A `sum` /
* `avg` over the exact-decimal column (`numeric(65,30)` / `DECIMAL(65,30)`)
* whose value needs more than a double's ~15-17 significant digits, or an
* integer at or above 2^53, is ROUNDED to the nearest double — declared, not
* silent: it is the same bound `formatOutput` already puts on a `find()` read of
* that column (#16318, `valueSchemaFor`'s `z.number().finite()`, ADR-0104 D1),
* and the bound the rows path has always had (`toNumber` sums JS doubles). A
* value-dependent type — a number when it fits, a string when it does not — was
* rejected: it would reopen this defect for exactly the large totals, where a
* `having` `$in` or a chart silently stops matching. Only a string `Number()`
* reads as NaN (PostgreSQL's `numeric` `'NaN'`) is left as written, the
* presenter's existing rule.
*
* A `Record` over `AggregationFunction` on purpose: a function that joins the
* declared vocabulary without an answer here fails `tsc` rather than reaching a
* caller unpresented.
*/
const AGGREGATE_ANSWER_KIND: Readonly<Record<AggregationFunction, 'number' | 'column'>> = {
count: 'number',
count_distinct: 'number',
sum: 'number',
avg: 'number',
min: 'column',
max: 'column',
};
// [#20335, #20889] `AGGREGATE_ANSWER_KIND` — what each declared aggregate
// function ANSWERS, and the one-double precision policy for that answer — lives
// in `@objectstack/core` (`utils/aggregate-answer.ts`), imported above with its
// `'number'` presenter, so the analytics native-SQL face presents a count or a
// total with this driver's own rule.

/**
* [#20387] What each declared aggregate function ACCUMULATES IN on PostgreSQL
Expand Down Expand Up @@ -15839,16 +15791,11 @@ export class SqlDriver implements IDataDriver {
return presentAuditTimestampOutput(value);
case 'boolean':
return Boolean(value);
case 'number': {
// Only strings are repaired, exactly as in `formatOutput`: a fresh
// REAL/INTEGER column already yields a number, and genuinely
// non-numeric legacy junk is left intact rather than turned into NaN.
if (typeof value === 'string' && value.trim() !== '') {
const n = Number(value);
if (!Number.isNaN(n)) return n;
}
return value;
}
case 'number':
// [#20889] The presenter lives in `@objectstack/core`, beside
// `AGGREGATE_ANSWER_KIND`, so the analytics native-SQL face presents
// a count or a total with this same function.
return presentAsNumber(value);
}
}

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -242,16 +242,17 @@ for (const cell of CELLS) {
const before = { ...reads };
const res = await query({ measures: ['row_count'], dimensions: ['title_dim'] });
expect(res.status, JSON.stringify(res.body)).toBe(200);
const groups = (res.body.rows as Array<{ title_dim: string; row_count: number | string }>)
.map((r) => [r.title_dim, Number(r.row_count)] as const)
// [#20889] `row_count` is a JSON number on every dialect — never `"2"`.
const groups = (res.body.rows as Array<{ title_dim: string; row_count: unknown }>)
.map((r) => [r.title_dim, r.row_count] as const)
.sort(([a], [b]) => a.localeCompare(b));
expect(groups).toEqual([['x', 2], ['y', 1]]);
expect(reads.rawSql - before.rawSql, 'the native strategy answered').toBeGreaterThanOrEqual(1);

const joined = await query({ measures: ['row_count'], dimensions: ['acct_name'] });
expect(joined.status, JSON.stringify(joined.body)).toBe(200);
const joinedGroups = (joined.body.rows as Array<{ acct_name: string; row_count: number | string }>)
.map((r) => [r.acct_name, Number(r.row_count)] as const)
const joinedGroups = (joined.body.rows as Array<{ acct_name: string; row_count: unknown }>)
.map((r) => [r.acct_name, r.row_count] as const)
.sort(([a], [b]) => a.localeCompare(b));
expect(joinedGroups).toEqual([['A', 2], ['B', 1]]);
});
Expand Down
Loading
Loading