Revision history for Stats-LikeR 0.3214 2026-10-05 CDT [read_table(): delimited text] - colClasses read past the end of a field the parser had not copied: Atof() ran on into the separator and the next field whenever the separator could continue a number, so with sep '.' the field "1" of "1.2" came back as 1.2, and with sep 'E' the "1" of "1E2" as 100. The number is now parsed from a copy of the field alone. - A quoted first field starting with the comment marker lost it: a header written '"#a",b' came back as a column named a, and where the next row was text as well the header was dropped and that row taken in its place. A comment marker inside quotes is now text, as R's read.table reads it. - A sep regex under Unicode rules -- /u, which `use v5.12` and later put on every qr//, or one using \p{} -- was matched against each line's bytes, and took the trailing 0xA0 or 0x85 byte of a UTF-8 character for a no-break space or NEL: qr/\s+/u cut "voilà" into "voil\xC3" and a separator, and Cyrillic "Р" likewise, and never split on a U+2003. A valid UTF-8 line is now matched as UTF-8 by such a pattern, and by one under /a; a Latin-1 line is matched as bytes, as before, and the fields are the file's bytes either way. - A read error was reported with libc's strerror() of whatever errno held, which on a threaded perl is not the function $! calls, and which could be stale. The reason now comes from $!, and only when the failing read set one. [read_table(): Excel workbooks] - A phonetic guide () is no longer read as part of the text. Japanese Excel records the furigana for every string typed through the input method, and 漢字 read back as 漢字カンジ. - The shared-string table is found through the workbook's relationships instead of by its usual name. A workbook that stored it as xl/SharedStrings.xml read every string cell as empty, and the first data row silently became the header. - .xlsm, .xltx and .xltm files are read as workbooks. They were read as text and reported as an alignment error. - Attributes quoted with ' are read, as XML allows; a workbook written that way lost every sheet name and its sheets were matched by position. A sheet with no name no longer takes the key of a sheet really called SheetN. - The XML is read as an XML parser reads it: a commented-out row is not data, a CDATA section is its literal text, and a CR LF or lone CR in a cell is LF ( is still a CR). A character reference to something XML does not allow (�, a surrogate, anything past U+10FFFF) is left in the text, where it used to come back as bytes that are not UTF-8; and one to a noncharacter such as 􏿿 no longer dies on perls before 5.14. - Excel's _xHHHH_ escape for a character XML cannot hold (_x0000_ for NUL, _x005F_ for a literal underscore) is decoded in string cells, as Excel and LibreOffice decode it, so a cell with a control character in it reads back as itself; a formula's cached result is left as written. - An .xlsx read as an aoh leaked a row hash when colClasses refused a cell. Every cell is now converted before the row's hash is made. [read_table(): speed and memory] - The parser tests each field against na_strings as it cuts it, so an NA cell never gets a string buffer and the rows need no second pass. A 300,000 x 10 CSV with half its cells "." reads 25-30% faster (aoa 0.218 s to 0.153 s, aoh 0.285 s to 0.215 s, hoa 0.208 s to 0.147 s); one with no NA cells, 3-6% faster. - sep => qr/\s+/ is split in C rather than by a regex match per field, with the same rows, errors and warnings. A 300,000 x 10 whitespace- separated file reads as an aoa in 0.18 s rather than 0.35 s, and as an aoh in 0.24 s rather than 0.41 s. - An .xlsx's shared strings are kept once and shared by every cell that uses them: a 500,000-row sheet of categorical columns takes 188 MB instead of 249 MB, and reads 7% faster. On perls before 5.18 such a cell reports itself read-only to Scalar::Util::readonly, though assigning to it works as before. [write_table(): records read_table could not see] - A record whose only field was empty, undef or blank was written as a blank line, which read_table skips, so a one-column table lost those rows and every later row moved up. It is now written quoted, '""' or '" "', as csv.writer writes [''] (CPython 3.14.2 test_csv.py). A first field starting with '#' is quoted too, since read_table took the record for a comment. - An undef in an AoA's col_names was skipped, so every later name moved one column left: ['a', undef, 'c'] over [1, 2, 3] wrote "a,c,". It is now an empty header cell in its place. A reference in a header cell is refused as one in a data cell is, where it was written as its address. [write_table(): arguments] - sep => '', and a sep holding a NUL, a quote, a CR or a LF, are refused, as is a file name holding a NUL, which used to be cut short there. - row_names => 'name' now heads an AoA's or a flat hash's 1..n label column, as it heads a HoH's, where it was ignored. A HoA row_names that names no column dies, and AoH rows without it draw one warning with their count, where both gave empty labels silently. - Rows that are restricted hashes missing a column no longer die with "Attempt to access disallowed key". [write_table(): .xlsx workbooks Excel and openpyxl can open] - U+FFFE, U+FFFF and surrogates were written raw, and openpyxl refused the workbook as not well-formed. They, and the control characters that used to be dropped, are now written as Excel's _xHHHH_ escape, as XlsxWriter writes them. Every underscore that would open an escape is itself escaped as _x005F_, overlapping ones included, where XlsxWriter escapes only non-overlapping matches and a literal "_x005F_x0041_" came back from a left-to-right decoder as "_x005FA". - Excel's limits are enforced: 16384 columns, 1048576 rows with the header, 32767 characters in a cell. A 16385th column used to be written as XFE. xlsx_freeze_rows and xlsx_freeze_cols are bounded, where 2**32 + 1 froze one row. - A string becomes a number cell only when Excel would show it as it stands, so "007", "+5" and 20-digit IDs stay text, and "1e400" is no longer an infinite number. A value perl holds as a number is always a number cell, except Inf and NaN. [write_table(): LaTeX] - Every line of tex_comment is now a comment: a newline in it used to end the comment, and the rest was typeset. An object passed as tex_comment or xlsx_comment is stringified rather than dropped. - '<' becomes \textless{}, as '>' became \textgreater{}; it printed as an inverted exclamation mark. \ $ { } ^ ~ stay raw so cells can hold math and macros, and the documentation now says so. - A croak after the .tex file was opened no longer leaves its handle open. [write_table(): speed and memory] - Finding an AoH's or a HoH's columns no longer leaves an iterator on every row hash: 200,000 rows came back 22.9 MB heavier. The AoH write is 16% faster (0.342 s to 0.288 s) and the HoH write 13% faster (0.806 s to 0.704 s), and an AoH written as .xlsx peaks 0.7 MB over its data rather than 30 MB. 0.3213 2026-10-01 CDT [Every name is spelled with _ instead of .] - Every option name and result key that followed R in spelling itself with a dot now uses an underscore: p_value, conf_int, conf_level, statistic_name, null_value, fitted_values, df_residual, r_squared, row_names, col_names, output_type, na_strings, undef_val, Res_Df, CI_lower, SE_theta and the rest, 112 names in all. Perl parses $r->{p.value} as "p" . "value", so every dotted key had to be quoted; $r->{p_value} needs no quotes. - The dotted spellings are gone rather than kept as aliases: conf.level, output.type, row.names and the rest are now refused as unknown arguments, where many functions used to accept both. density()'s give.Rkern becomes give_rkern, the underscore spelling it already took. - rank(), cohen_d(), age_standardize() and TukeyHSD() never checked their option names, so an old dotted name would have been ignored without a word, and rank() would have ranked 'ties.method' as data. All four now refuse names they do not know. - Only names change. Option values keep R's spelling (alternative => 'two.sided', type => 'one.sample', family => 'negative.binomial'), and so do names built from the data: merge()'s default suffixes .x and .y, pivot_table()'s sep, the .1 and .2 that make a repeated name unique, read_table()'s exploded VCF columns such as NA00001.GT, and R's table headers Std. Error and Resid. Df. [assign(): consistent context, shape detection, no half-applied calls] - A per-row coderef was called in list context for the first row, to tell a whole-column return apart, and in scalar context for every other row, so anything that answers differently in the two gave the first row a different kind of value: sub { $_->{s} =~ /(\d+)/ } stored the digits in row 0 and 1 in every other row, and a failed match was undef in row 0 and '' elsewhere. Every row is now called in list context; a later row that returns more than one value dies instead of storing a count. - A hash frame was classified as HoA or HoH by whichever value values() returned first, so a HoA holding a scalar or undef entry worked or died from one run to the next with hash order (167 of 200 runs died on a frame with five such entries). Every value is now looked at; a hash of both arrays and hashes, or of plain scalars only, dies with a message that says so. - Rows, value types, arrayref lengths and map_cell targets are all checked before the first write, so a bad last row or a bad later pair no longer leaves the earlier rows or pairs applied. - An empty hash takes its row count from its first arrayref value, so assign({}, x => [1, 2, 3]) makes a column where it used to die. - On a HoA the row loop is now C (_hoa_assign), and the row view aliases the frame: one hash for the whole call, whose values are the current row's own cells rather than copies. A write through $_->{col} therefore reaches the frame, as it always has for an AoH or HoH, and map_cell's $_ is the cell itself. Building a fresh hash of copies per row was 73% of the call, and C alone would have cut that by only a quarter. A per-row coderef on 200000 rows x 16 columns went from 0.41s to 0.062s, and on 1000000 x 64 from 6.6s to 0.94s; map_cell from 0.23s to 0.055s and from 4.1s to 0.93s. A key the block adds to the view is gone on the next row, and a block that keeps $_ keeps that one view. A tied frame or column keeps the perl loop and its copies. - A whole-column result on a HoA is installed without a second copy of the list, which takes 1.5 MB off the call's peak on 200000 rows. [vals(), avals(): a column no row has is an error] - On an AoH or HoH, asking for a column that no row has returned one undef per row, so a misspelt name went unnoticed until something downstream choked on the undefs. Both now die with 'vals: no column named "Method"' (or 'avals: ...'), as a HoA already did, and a HoA's missing column says the same thing rather than "not found or is not an array-ref". A column that some rows lack, or that holds only undef, is still read as before, and an empty frame still gives an empty result. [read_table(): a comment as wide as the header] - A leading comment that split into as many fields as the next line was taken for a commented-out header, and the file's real header then read as the first data row, with no warning: "# written by foo, v2" before "id,val" named the columns "written by foo" and " v2". A commented line is now the header only when the line after it looks like data -- a number or an na_strings token in one of its fields, or nothing but empty fields; otherwise it is a comment, as in R and pandas. Over rows of text alone a commented-out header cannot be told from a comment, and the first row is read as the header. [read_table(): VCF files] - A VCF could not be read. Its "##" meta lines have the marker hugging the text ("##fileformat=VCFv4.2"), which the parser passes on as a possible commented-out header, and the first such line was taken for the header on the spot: the file was one column named "fileformat=VCFv4.2", and the "#CHROM" line an alignment error. That happened with comment => '##' and with '#' alike. A run of these lines before the header is now a run of candidates; the last, the one next to the data, is tried under the rule above, and if it fails the next line is the header. R's tests/reg-IO2.R test.dat ("#comment", "#another", "#", then a C1 C2 C3 header) now reads with header = TRUE as R reads it, where it was an alignment error. - A file named .vcf, .vcf.gz or .vcf.bgz is read with a tab sep, a "##" comment marker and the "#" taken off "#CHROM", so read_table('x.vcf.gz') needs no options. The cases are htslib 1.21's own test VCFs, with bcftools 1.21's reading of them as the expected values. - Each sample column of a VCF is split on ':' into one column per FORMAT key, named ".", and the records come back as a hoh keyed by CHROM:POS:REF:ALT. FORMAT may change from record to record, so the keys are those of every record read; a key a record lacks, or a value a sample drops off the end, is undef. Values stay text. 'output_type' still gives the other shapes, 'row_names' another key, and the new explode => 0 the file's own columns. A filter sees the file's own columns, since it runs while the file is read. The cases are bcftools 1.21's test VCFs, with each value as bcftools query prints it. The split is done in C: on a 3,499,678-record single-sample .vcf.gz it adds 2.6 s to the 9.1 s a plain hoh takes, where the first version, in perl, added 12.6 s. [read_table(): CR line ends] - A file whose lines end in a bare CR, as classic Mac OS wrote them, was read as one line with every CR dropped, and came back as []. One whose start has a CR and no LF is now split on CR, as R's scan() and pandas' C tokenizer split it. A stray CR in an LF or CRLF file is still dropped. The cases are pandas 3.0.4's test_cr_delimited and test_tokenize_CR_with_quoting, and R's PR#2469. [read_table(): repeated row names] - A hoh warned once for every row that repeated an earlier row's name: 40,820 warnings for a 300,000-row file of random names. It now warns once for the file, with the count and the first repeated name and row. A single repeat keeps the words it has always had. [read_table(): speed] - The CSV parser stores into its own arrays without going through av_push(), finds the next special byte with one table lookup instead of three tests, hands its row buffer over as an aoa's output row instead of copying it, and looks a hoh's row name up once instead of twice. On a 300,000 x 10 CSV, best of seven: aoa 0.147 s to 0.117 s, hoa 0.204 s to 0.178 s, aoh 0.179 s to 0.165 s, and a 300,000-row hoh 0.357 s to 0.339 s. [read_table(): fixes from a review] - VCF explode aborted ("realloc(): invalid pointer") when two output columns had one name and the shape was hoa: a sample name repeated in the header, as merged files have, or sample "X" with key "Y.Z" beside sample "X.Y" with key "Z". The second column's array replaced the first's while rows were still pushed into the first. An aoh or hoh kept the later column without a word. Every shape but aoa now keeps the later column and warns as the plain read does about a repeated name. - A commented-out header ("# id,val") was cut by a perl split(), which ignored quoting and cut before the marker came off: '# "x,y",z' gave three names and "#ab" gave ('', 'a', 'b'). The widths then failed to match the data, so the first data row became the header and was lost, silently. The line is now cut by the parser itself. - In an .xlsx, a first cell starting with '#' was taken for a commented-out header, though 'comment' does not apply to an .xlsx: a column "#id" with values "#a1" came back as [], and "# of items" lost its "# ". - A filter field named by two keys, its number and its name, ran only one of the two, and which one followed hash order, so the rows kept changed between runs. Both now run, in key order. A key that is a column's name is that column before it is a field number, so a column named "2021" can be filtered on; 0 is still the whole row. Under a repeated column name, a filter on an earlier field now sees that field's own value, and its write-back no longer replaces the value the row keeps for the name. - 'sheet' takes a name before an index. Of sheets "2024" and "2023", neither could be asked for by the name the whole-book read returns it under, and with sheets "2" and "1", sheet => '1' returned "2". - An .xlsx whose elements carry a namespace prefix (, as the Open XML SDK writes them) read as [] without a warning; it is now read. A cell written open and empty (, ) past a row's last value no longer adds an unnamed column, as the self-closing already did not. - Text passed in as characters, under `use utf8` or decoded, is compared with the file as UTF-8 bytes. A qr// sep holding a character past 0xFF never matched, nor did such an na_strings token, filter key, row_names or sheet name, while a literal sep did. A UTF-8 qr// sep now matches each line as UTF-8 and refuses a line that is not. - row_names with an aoh or hoa, and sheet on anything but an .xlsx, were ignored; both are now errors, as row_names with an aoa already was. An xz, LZMA, zstd or lzop file is refused by name, from R's magic numbers, where it was read as text and reported as an alignment error. [t_test(): tied input] - A sample's length came from av_len(), which is FETCHSIZE on a tied array, but its elements were read straight out of the array's own storage, which a tied array does not use. A tied array that still held real elements from before the tie was read past their end and segfaulted; a freshly tied one read as empty and croaked "'x' needs at least 2 elements". Elements are now read through FETCH on a tied array. - Get magic was not run before testing whether an element or argument was defined, so a tied element read as missing and was dropped, and a tied scalar holding the array reference was refused with "Usage: ...". Both now work, and each element is fetched once. - A sample's buffer was sized from one call to FETCHSIZE and filled up to a second. A tied array whose FETCHSIZE grew in between was written past the end of the buffer and segfaulted. The length the buffer was sized from is now the one read. [t_test(): far-tail p-values] - The p-value squared the t statistic itself, so for |t| above sqrt(DBL_MAX), about 1.34e154, it came back as exactly 0 where R's is a positive number. That tail is still far above underflow at df <= 2: t_test([1e-160, 2e-160], mu => 1) gave 0 for R's 3.18e-161. It now goes through the same tail code as pt(), which has R's asymptotic branch. R's own check for this, pt(-a, df = 1) == pcauchy(-a) out to a = 1e300 (tests/d-p-q-r-tst-2.R), is now in t/t_test.tails.R.t. [t_test(): the variance] - Welford's update was replaced with R's two-pass variance. It lost digits on data with a large offset: four values 1e-4 apart at 1e10 gave t = 1.54414e14 for an exact 1.55003e14, and a 1e6 +- 1e-6 sample's standard error was 6e-5 out against R's 7e-9. - It goes past R in two ways. The squared deviations have (sum of deviations)^2 / n subtracted, the Chan-Golub-LeVeque correction, and the same sum is carried into t as the part of the mean no double holds. On a sample whose spread is a few dozen ulps of its mean, R's t is out in the third digit and this one is within 2e-16 of exact rational arithmetic. - The data are scaled by a power of two before squaring, so squares past DBL_MAX no longer overflow. t_test([1e154, -1e154, 3e153], [2e154, -1e153, 5e153]) returned t = -0 and df = NaN. Its Welch df and standard error are now formed from ratios and are finite, where R's df, written in stderr^4, is NaN. [t_test(): arguments] - A NaN conf_level passed the range check and gave an interval of (-Inf, Inf); a NaN mu gave a NaN t. Both now croak, as R stops on both. - An undef mu or conf_level, R's NULL, was read as 0, and a reference was read as its address. Both croak "'mu' must be a single number", as R does; an object that overloads numification is still accepted. - The result carries stderr, the standard error the statistic divides by, which R's t.test() has returned since 3.6.0. - conf_level 0 and 1 were refused; R accepts both, as the ends of its range. 1 now gives (-Inf, Inf) and 0 a point at the estimate, as in R. This also stops a conf_level within an ulp of 1, such as 1 - 1e-20, which is 1.0 on a double, from croaking. [t_test(): memory] - A two-sample test copied x and y into separate buffers at once. x is now reduced to its moments before y is read into the same buffer, so the extra memory peaks at max(nx, ny) values instead of nx + ny: 800 MB less on a double build for two groups of 1e8. [qt(): Newton instead of bisection] - qt_tail(), the t quantile behind qt(), t_test()'s confidence interval, power_t_test(), svyglm(), ivreg() and lmer(), bisected all the way to adjacent doubles. That took 40 to 55 evaluations of the t tail, about twice as many on quadmath, and was most of t_test()'s time on samples under about a thousand values. It is now Newton on the log tail against log t, kept inside the same bracket and falling back to the midpoint when a step would leave it. qt(0.975, df) takes 2 to 9 us instead of 11 to 33. Results agree with the old ones to 7e-15 relative except near p = 0.5, where both are limited by the tail's own rounding and, where checked against mpmath, the new ones are the closer. [Incomplete beta: a NaN argument hung __float128 builds] - The continued fraction behind pt(), pf() and the binomial tails cast its iteration cap to long, and a NaN shape parameter reached the cast. That is undefined behaviour. A double build got LONG_MIN and skipped the loop; a quadmath build got a count it never finished, so t_test() on a sample holding an infinity, whose Welch df is NaN as R's is, hung. incbeta_xy() now returns NaN for a NaN argument, and the cap check sends a NaN to the ceiling instead of the cast. [Tests: every t.test() call in R] - t/t_test.R.scipy.t now carries every t.test() call in R's sources and tests: the examples of ?t.test, ?sleep, ?ks.test (drawn under R CMD check's set.seed(1)), ?array2DF and ?pairwise.t.test, R-intro, the tcltk demo, and reg-tests-1a, -1e and -2. Each is crossed over every alternative, var_equal, mu and conf_level, and the examples are also checked against R's pinned .Rout.save output at the digits R printed. PR#18901's stopifnot() is reproduced in full, where the test had only checked that the call did not crash. - The documentation said x needs 2 values even under var_equal, where either sample may have 1. It now says so. [logrank_test(): a group with no expected events] - A group never at risk at an event time -- one whose subjects are all censored before the first event -- has no expected events and a zero row and column in the variance matrix. Every group was kept in the test and the last one dropped, so that matrix was singular, the solve failed, and the statistic was left at 0 with no warning: p = 1, on a df that counted the empty group. survival's own tests/difftest.R adds such a group of seven early censorings to aml and expects aml's answer back, 3.396 on 1 df, p = 0.065; this gave 0 on 2 df, p = 1. Such groups are now left out of the test, as survdiff() leaves them out. With only one group left the statistic is 0 on 0 df with p = 1, as in survdiff(); with no events at all it is NaN on 0 df, where survdiff() gives df = -1. A variance matrix that is still singular -- every event time empties the risk set -- croaks, as survdiff() stops, instead of reporting p = 1. - An event time with one subject at risk was skipped outright, to keep the variance term's n - 1 divisor away from zero, so the last subject dying went uncounted in observed and expected alike. The statistic was right, the two cancelling, but the group's observed and expected events were each one too few. Only the variance term is skipped now. [survfit(): the median on a flat stretch at 0.5] - When S(t) steps onto exactly 0.5, survival:::survmean() puts the median halfway between that time and the next drop. survfit() took the start of the flat stretch instead, so 4 deaths among 8 at t = 5 with the next at t = 6 gave a median of 5 where R gives 5.5. It now follows R's rule, to R's tolerance of 2^-26 for "is 0.5". [Tests: survival] - t/survival.R.t pins logrank_test() and survfit() to R 4.6.1 with survival 3.8-12 at full precision, on aml and aml3 from tests/difftest.R, test1 from tests/quantile.R, and data built to reach each new branch. t/survival.t had covered aml alone. The generator is t/survival.R.R. [agg(): the split moves to XS] - Grouping is now done in C: one pass hashes each row's `by` cells into its group, a second drops each aggregated cell straight into its group's array. The perl split copied every needed column, built a row index per group and sliced a second copy out of the columns for each. On 1e6 AoH rows in 1000 groups agg() goes from 0.59 s to 0.14 s, and from 157 MB above the frame to 31 MB; ungrouped, from 0.56 s and 162 MB to 0.10 s and 30 MB. Plain-number cells that only the numeric reducers read are shared, not copied, and nothing is written back to the caller's scalars: keying on a numeric column no longer leaves a cached string on every cell, as "v$v" did. [agg(): fixes] - A coderef aggregator ran in list context, so one returning an empty list or several values shifted every later column of the row: { a => sub { grep { .. } @{ $_[0] } }, b => 'sum' } put b's sum under a. It now runs in scalar context. - Output names collided silently. Aggregating a `by` column (by => 'g', agg => { g => 'count' }) overwrote the group's key, and in hoa output left that column twice as long as the others; two coderefs on one column were both "fn", so only the second survived. A `by` column now aggregates to "_", several coderefs are fn1, fn2, .., and any name still generated twice dies, except under output_type aoa, whose columns are positional. - Groups sorted wrong whenever any key was undef or any `by` column held strings: one numeric-or-string decision covered every key column, and an undef made it "string", so 10 sorted before 2. Each column is now compared on its own terms, undef last and NaN after the numbers. A NaN key used to die inside sort under the module's FATAL warnings. pivot_table() shares the sort, and gets the same fix. - Keys with several columns were joined with a bare "\x1e", so ("p\x1evq", "r") and ("p", "q\x1evr") were one group. Cells are now length-prefixed, as drop_duplicates() and merge() key them. - skipna => 0 did not reach min and max, which pandas' skipna=False does. It now covers mean, median, sum, sd, var, min, max and mode. - A column in no row -- a misspelled name -- came back as a column of undef; it now dies, as pandas' KeyError does. An undef `by` column, a non-integer AoA position, a non-array HoA column and a defined row that is not a reference now die with a message saying so, rather than on a perl warning. A non-numeric cell now names the reducer, column and group. An unknown aggregator is caught before any group is reduced, so an empty frame no longer lets one through. - An empty frame aggregated without `by` gave no rows; it is now one row, as pandas' df.agg gives, with count and n 0. [Tests: agg()] - t/agg.R.pandas.t pins agg() to R 4.6.1's aggregate() -- every aggregate.Rd example that is a data frame, and the reg-tests-1a/1c/1d cases including PR#15004's 21 grouping columns -- and to pandas 3.0.4's own groupby tests, among them every float64 row of GH#15675's skipna parametrisations. Every case runs through all four input shapes. The generators are t/agg.R.pandas.R and t/agg.R.pandas.py. [anova(): the model R fits] - anova(\%data, formula) had a formula parser of its own, and it fitted models R does not. Every factor in an interaction was coded by contrasts, ignoring R's margin rule: y ~ a + a:b (b nested in a) gave a:b 2 df on warpbreaks where R gives 4, y ~ a:b gave 2 where R gives 5, and y ~ g + g:x did not fit a slope per group. Because that coding depended on which level came first, y ~ a:b gave a sum of squares of 5.40 or 22.08 on the same rows in a different order. Terms were taken in formula order rather than R's (main effects, then two-way interactions, ..), so y ~ a + b:a + b attributed b after the interaction. y ~ 0 + x was fitted with an intercept; y ~ x - 1 read "x - 1" as a column name and died on "fewer than 2 complete observations"; y ~ ., offset(), y ~ 1 and a*b:x with b a factor all died the same way or another; and y ~ a:b + b:a fitted the interaction twice. anova() now reads the formula with lm()'s parser and builds the design with lm()'s, and all of these match R 4.6.1. - A column is aliased by R's rule, when what the earlier columns leave of it has a norm below 1e-7 of its own (lm.fit's tol, dqrdc2.f). The test was 1e-10 on the largest absolute entry. - Data is read as lm() reads it, so a hash of hashes is accepted and a hash of arrays whose columns differ in length dies, rather than fitting on the response's length. [anova(): model comparison] - A list of formulas whose responses differ was compared on RSSs of different responses: 'y ~ a' against 'log(y) ~ a + b' gave F = 91.4. As anova.lmlist() does, a model whose response is not the first model's is now dropped with a warning, and if one model is left its own table is returned. - F follows stat.anova(). A step with Df > 0 and a negative Sum of Sq (models that are not nested) gave F < 0 and a p-value of 1; F and Pr(>F) are now left out, R's NA. A step with Df < 0 -- models listed largest first, as anova.lm.Rd's "unconventional order" example does -- had no F; it is now tested on abs(Df), as R tests it. [anova(): speed and memory] - The fit held the whole n-by-p design, one allocation per row, plus a dense copy of every factor's indicator columns, and reduced it by Householder QR with loops that walked each column down those rows. It now rotates one row at a time into a p-by-p triangle (Gentleman's Givens rotations, AS 274, as R's biglm does) and keeps nothing that grows with n. y ~ g*h + x on 1e5 rows with a 50-level g (201 design columns) went from 47.5 s and 196 MB to 1.6 s and under 1 MB (R's anova(lm()) takes 3.1 s); with a 200-level g (801 columns), from 638 s and 768 MB to 23 s and 2 MB. - Factor levels are found with a hash rather than a linear search per row. This is lm()'s design builder, so lm(), glm() and every model built on it gain it too. [lm(), glm(): a:b and b:a] - The term list dropped a repeated term only when it was spelled the same way, so y ~ a:b + b:a carried the interaction twice, the second copy aliased. a:b and b:a are now one term, as in R. [Tests: anova()] - t/anova.R.t pins anova() to R 4.6.1's anova.lm() and anova.lmlist(): anova.lm.Rd's LifeCycleSavings tables, including the unconventional order that stats-Ex.Rout.save pins; warpbreaks.Rd's model and the same data nested, alone, out of order and without an intercept; npk.Rd's confounded design; ToothGrowth with a slope per supplement; reg-tests-2.R's offset (PR#8049) and 0-rank models. The generator is t/anova.R.R. t/anova.t gains leak checks for the paths that croak after the designs are built. [aov(): the model R fits] - aov() had a formula parser and a Householder QR of its own. It now reads the formula with lm()'s parser and fits it as anova() does, which fixes the models below. - Groups whose names look like numbers were fitted as one slope. aov({1 => [..], 2 => [..], 3 => [..]}), the documented no-formula form, gave Group 1 df where R's stack() makes a three-level factor with 2: F = 26.43 for R's 12.04 on the same data. The stacked Group is now a factor whatever its labels are. - a*b*c crossed only its first `*`, and offset() was read as a column name, so both died with a misleading "0 degrees of freedom". - Every factor in an interaction was coded by contrasts, so y ~ a:b, y ~ a + a:b and y ~ g + g:x died with "requires its main effects", and y ~ g - 1 lost a level of g. Factors are now coded by R's margin rule. - Terms were taken in the order written, so y ~ a + b:a + b died the same way and a:b + b:a was two terms. They are now in R's order, by degree, and named as R names them. - A column was aliased when its largest remaining entry fell below 1e-10 of its largest entry. It is now R's lm.fit rule: aliased when the part the earlier columns leave has a norm below 1e-7 of its own. - The table now follows anova(): a 0-df term has no Mean Sq, and no term has an F test when there are no residual df or the fit is exact, where they used to be 0 and NaN. Two complete rows are enough, as in R; "0 degrees of freedom" used to stop any fit with no residual df. - fitted_values include the offsets, as R's do, and are keyed by a row_names column when there is one, as lm()'s are. [aov(): group_stats] - With a formula, group_stats held the mean and count of every column of the data, over every row: the response's overall mean, and NaN for each factor. It now holds the response's mean and count in each level of each factor of the model, over the rows fitted. With one factor it is keyed by level, as the stacked form and oneway_test() key it; with several, by factor and then level. The stacked form's values are unchanged. [aov(): speed and memory] - The design was held twice, n by p each, and every level of every factor in every row was a string copy. The fit now keeps a p-by-p triangle and re-reads the rows for fitted_values. y ~ g*h + x over 1e5 rows with a 50-level g went from 40.9 s and 367 MB to 1.8 s and 52 MB; anova() takes 1.6 s. [lm(), glm(), anova(): interaction coefficient names] - An interaction's coefficients were named in the order its variables were written, so y ~ w + t + t:w named them tL:wB where R has wB:tL. A term's variables now follow their first appearance in the formula, as R's terms() orders them, which also gives the columns of the term R's order. [Tests: aov()] - t/aov.R.t pins aov()'s table, coefficients, fitted values and group_stats to R 4.6.1: aov.Rd's npk models, warpbreaks.Rd, reg-tests-1b's warpbreaks with NA rows, reg-tests-3's oats (PR#7829), ToothGrowth with a slope per supplement and with an offset, reg-tests-2's PR#16437 with and without an intercept, and PlantGrowth stacked, under its own group names and under "1", "10" and "2". The generator is t/aov.R.R. One divergence is recorded there: R drops a factor level whose rows were all dropped for NA, and lm()'s design builder keeps it as an aliased column. [each() survives every XS function] - Every XS function that walked a hash it was handed reset that hash's iterator, the one each(), keys() and values() share, so a caller part way through while (my ($k, $v) = each %$df) started again and saw keys twice: filter, csort, col2col, lm, glm, predict, aov, merge, write_table, value_counts, group_by, p_adjust, prcomp, transpose, the hoa2*/hoh2*/aoh2* converters and the rest, 40 functions in all, on the table, on its rows, and on other hash arguments such as a model's coefficients. Each walk now saves the iterator first and puts it back afterwards, including when it croaks part-way. agg, select_cols, drop_cols, rename_cols, drop_duplicates and assign still reset it: their perl side calls keys() itself, as perl's own keys() always does. t/iter.keep.t checks each function. [Crashes] - predict() segfaulted on HoH newdata with a row that was not a hash ref: the shape came from whichever row the hash handed over first, and every other value was dereferenced as a hash, so the crash came and went with hash order. Every row is now checked, and the walk is bounded by its buffers, which a tied hash (reporting no keys) would have overrun. - csort(), prcomp() and sample() segfaulted on a locked hash (Hash::Util) with a deleted key. Each sized its work by a count that includes the placeholder a delete leaves behind, so the walk filled one slot fewer than the loops after it read. Each now uses the count the walk saw. [csort(): HoH input is no longer quadratic] - Folding a HoH into rows sorted the row names with an insertion sort: 82 s for 200000 rows, all of it there. It is now qsort() with the same bytewise order, 0.61 s, and the output is unchanged (checked against 0.3212 in all three output shapes, under three hash seeds, with tied sort keys and UTF-8 row names). [Leaks on croak] - hoh2hoa leaked its result and scratch tables on a non-hash row or a row_names collision; kruskal_test leaked its observations and labels when reading a label died (38 KB over 200 calls under valgrind); col2col leaked its row table when a tied row could not be walked; and prcomp freed by hand only on its own croaks. All four now leave their tables to perl -- mortal, or on the save stack -- so nothing leaks however the call ends. t/croak.leaks.t covers each. [Tied hashes] - A tied hash -- the frame itself, a row of a HoH or AoH, or another hash argument -- keeps nothing where the XS looked: an iterated entry has no value until it is fetched, a fetched value is a placeholder until its get magic runs, two fetches share one entry, and every key count reads 0. csort, value_counts, kruskal_test, merge, agg and drop_duplicates segfaulted on one; lm, glm, predict, aov, oneway_test, p_adjust, chisq_test, fisher_test, sample, group_by, col2col, hoa2aoh, hoa2hoh, hoh2hoa and the HoH paths of select_cols, drop_cols and rename_cols refused it as empty or malformed; vals and avals returned an empty list without a word, and ljoin and add_data joined nothing from it or into it. Each now gives the answer it gives for the same data untied, which t/tied.hashes.t checks for every one of them. - merge() takes a HoH frame's rows in row-name order rather than hash order, so the same merge gives the same rows in the same order on every run, and the same as a tied copy of the frame. Hash order differs between runs, so nothing could have depended on it; on 200000 rows the sorted merge also ran faster (0.18 s against 0.36 s). - group_by()'s filter left an extra reference on every filtered value, plain data included, so those values could never be freed. value_counts() leaked its result when reading a value died. - chisq_test() and fisher_test() reported a key missing from a tied row, or from a tied 'p', as an undef cell ("cell {r2}{b} is undef"), because a tied hash's fetch hands back an entry whether or not the key is there. They now ask first and say the key is missing, as they do for a plain hash. - A tied value was FETCHed more than once per call by most of the frame functions: every pass over a tied frame (the shape probe, a key count, a key union, the read itself) fetched each row or column again, and a tied column read once to match or test on and again to copy out had every cell fetched twice. lm() and glm() fetched a tied frame's column once per row of the fit; select_cols, drop_cols, rename_cols, hoh2hoa and group_by fetched each row of a tied HoH three times. The answers were right, but a tie whose FETCH does real work paid for each one. A function that reads a tied frame now takes a plain copy of it first, one FETCH per value, and works on that; where it builds its result afresh it copies the frame's tied rows or columns too. Nothing is copied for a frame with nothing tied in it. rename_cols() in void context still renames the tied frame itself. t/tied.fetch.once.t counts the FETCHes of 50 calls. [Tied arrays] - sample() on a tied array segfaulted: it read AvARRAY(), which a tied array leaves empty. It now fetches each drawn element, and with the same seed draws what it draws from a plain copy of the array. - vals() and avals() returned one undef per row for a tied HoA column, and croaked "AoH row 0 is undef" on a tied AoH. col2col() croaked "no usable columns found" on a tied AoH. chisq_test() and fisher_test() croaked "cell [0][0] is undef" on an AoA whose rows are tied, and chisq_test() on a tied vector. Each read a tied array's storage directly, or tested a fetched value before its get magic had run; each now gives the untied answer, which t/tied.frames.t checks. - More of them tested a fetched cell or row before its get magic had run. group_by() returned {} for a HoA with tied columns. kruskal_test(), aov() and oneway_test() croaked on tied groups ("all groups must contain data", "fewer than 2 complete observations", "observation 0 is undefined or non-numeric"), and lm() and glm() on tied HoA columns ("0 degrees of freedom"). hoa2hoh() croaked on a tied key column, binom_test() on a tied vector ("successes is undef"), and epi_2x2(), cmh_test(), survfit(), logrank_test() and coxph() on tied vectors ("... at index 0 is undef"). fisher_test() refused a tied outer AoA with "each row must be an array ref". Each now copies the tied columns or rows first, or runs the cell's magic before testing it, and gives the untied answer; t/tied.frames.t checks them. [vals(), avals(): HoH input is no longer quadratic] - The row keys of a HoH were put in order with an insertion sort: 5.25 s for 50000 rows. It is now qsort() with the same sv_cmp() order, 0.05 s for vals and avals together, and the output is unchanged. [prcomp(): column names and tied input] - Column names were copied as C strings and looked up again by their strlen(), which lost a UTF-8 name's flag and cut a name at an embedded NUL. A HoA with a column named "\x{263a}" croaked "cannot be looked up by name", an AoH or HoH with one read every cell as missing and croaked "0 valid observations", and varnames came back without the flag. The names are now the keys themselves, as SVs, and come back intact; a name holding a NUL, which used to croak, now simply works. Their order is sv_cmp(), which get_all_columns() already uses, and is unchanged for ASCII names (checked against 0.3212 to 15 digits). - A tied frame hash read as empty, and a tied row, row hash or column cell was tested before its value was fetched, so it read as undef and the row was dropped without a word: an AoA with one tied row of four was decomposed on three. Every shape now gives the untied answer. [filter(): tied frames] - A tied HoA or HoH died as "hash data frame must be a hash of arrays (HoA) or a hash of hashes (HoH)", because the shape test read a value a tied hash only fetches on request, and a tied HoA would then have had no columns, its count read where a tied hash reports 0. Both work now. [cfilter(): crashes, leaks and memory] - A predicate that replaced a row or column of the data it was filtering, as in cfilter(\@aoh, keep => sub { $aoh[1] = 5; 1 }), segfaulted: the output was rebuilt from the caller's data on the assumption that it still had the shape checked before the predicate ran. It is now checked again, and such a call dies with the same message as bad input would. Predicate columns are built from the rows as they were when cfilter was called, so a predicate that edits the data cannot misalign a later column against the one named by against. - Every croak after the scratch tables were made leaked them: an unknown column name, an unknown against column, a bad AoH element, and a predicate that dies, which leaked a full copy of the table. They are now mortal and are freed by the croak. - Selecting from an AoH or HoH kept two short-lived SVs per cell alive until the caller's statement ended, one for each key it looked at. On a 20000 x 50 AoH, keeping one column raised the peak by 101 MB for a result that is 4.7 MB in pure perl; it is now 6.7 MB, and the call takes 0.072s instead of 0.111s. - A predicate no longer works on a second copy of the whole table made up front. Each column is copied just before its call and freed just after it, which halves the peak on a 20000 x 50 HoA (62 MB to 31 MB, the size of the result) and takes 196 MB to 72 MB on the same AoH. - A tied hash died with "hash values must be array refs", because the value was tested before the tie had fetched it. Tied hashes and arrays now work at every level: the table, a row, a column. - An each() loop the caller was part-way through, on the table or on any of its rows, started again after cfilter, because every walk of a hash reset the iterator that each() shares; a croak in mid-walk left the next each() starting part-way in. cfilter now saves the iterator and puts it back on the way out, croaks included. If the predicate deletes the entry that each() was on, perl has freed it, and the iterator is left reset as before rather than pointed at freed memory. - The documentation said the predicate is given only the defined cells. By default it is given every cell, undef included; na => 'omit' and against are what drop them, and both are now documented. [Strings outside ASCII: one string, one key] - Several functions used a string as a hash key, or compared two, by its bytes alone and dropped its UTF-8 flag, so the two spellings perl has for one string ("caf\xe9" and its upgraded copy) were two keys, two strings with one spelling ("caf\xc3\xa9" and the upgraded "caf\xe9") were one, and a name outside Latin-1 was looked up by its bytes and not found. Each now keys a string as a perl hash would, and t/utf8.keys.regressions.t has a case for every one. - filter() named a HoA column it built after its name's bytes and, looking the cells up by those bytes too, filled it with undef: AoH, HoH and HoA into HoA and HoA into AoH, with a code block or a col() expression. A HoH's row keys lost the flag the same way. - csort() said "column not found" for a column named outside ASCII, in an AoH or a HoA, and value_counts() counted nothing for one. - merge() would not join the two spellings of one key and did join two keys with one spelling; drop_duplicates() and mode() erred the same ways. group_by(), agg() and uniq() were already right, and merge() and drop_duplicates() now share their code. - survfit() and logrank_test() split a group label by spelling and returned one outside Latin-1, as a stratum name and in 'groups', as its bytes. A label also ended at its first NUL. coxph() split a strata => \@labels level by spelling. - csort() rethrew a comparator's die as a string, so an exception object came back as "HASH(0x...) at ..." and a UTF-8 message as its bytes. $@ itself is now rethrown. col2col() kept the flag off a callback's error. [read_table(): filters run in the parser] - A filter kept every row of a read in read_table's perl closure, which built each row's hash and called the subs from there. The parser now calls them itself, with the same $_, %_, arguments and write-back. On a 300,000 x 5 CSV a filtered read took 0.58 s to 1.07 s, by shape and filter, and takes 0.23 s to 0.37 s; 0.08 s to 0.18 s unfiltered. t/read_table.filter.t reads every case both ways and requires the same result, the same values seen by each filter and the same warnings. - A filter that is not a code reference is refused before the file is read, where it died on the first row that reached it. [read_table(): colClasses] - New option, R's colClasses: 'numeric' (or 'double', 'real') stores a column as numbers, 'integer' as integers, and 'character' or undef as read, given as a hash by name, a list by position (recycled when short) or one class for all. A field that is not a number of the kind declared is an error naming the column, the row and the text. What is accepted is what R's scan() accepts, except that an integer may be any IV, hex is refused and an exponent needs digits. On that CSV with three of five columns declared, a hoa took 74 MB instead of 116 MB, and the read 0.095 s instead of 0.086 s. The cases are R's PR#16478 and reg-IO2.R and pandas 3.0.4's test_dtypes_basic.py. [merge(): a hole in suffixes] - suffixes => \@s segfaulted when @s had a hole (my @s; $s[1] = '.y'). A hole is now the undef it prints as, as an undef element already was. [xs.check.pl] - Its output is one line per finding, and it exits 1 while any remain; each finding used to be a five-line stack trace, 196 of them in 1177 lines. XS::Check's type checks read one declaration per variable name for the whole file, so 41 of its 43 were about another function's variable; xs.check.pl now checks the declaration in scope. SvPV() calls whose encoding has been checked are listed in it, with the reason. 0.3212 2026-09-28 CDT [write_table(): a warning for a header cell with no name] - A hash of hashes writes its outer keys as a leading column by default, and unless row.names names that column its header cell is empty, so the file gives the column no name and every reader invents one: read_table calls it row_name, pandas "Unnamed: 0". write_table now warns when it writes that, and says that row.names => 'name' names the column and row.names => 0 drops it. This holds for LaTeX and .xlsx output too. - It also warns, for every shape, when a data column's header is empty: an empty or undef cell in an array of arrays' header row, or '' in col.names. The warning counts such columns and gives the first one's position in the file. - The empty label cell that row.names => 1 writes for the other shapes is not warned about. It is R's own layout, asked for explicitly, and row.names cannot name it for those shapes, so the warning could not be silenced. The file written is unchanged in every case, and quiet => 1 does not suppress the warnings, which are about the data and not the confirmation line. - The warnings are given after the file is written and closed and every buffer is released, so a __WARN__ handler that dies loses no output and leaks nothing. t/write_table.t covers each shape, each row.names mode, LaTeX and .xlsx, the dying handler, and leaks. [wilcox_test(): exact intervals on tied data, up to 330 times faster] - The exact interval on tied data is found as R finds it: at every shift between two neighbouring pairwise differences, up to m*n of them, the permutation distribution given the ranks is built and a tail read from it. It was rebuilt from scratch at every shift. It depends only on the sorted scores, which most shifts leave alone, so it is now kept and rebuilt only when they change; and the loop that builds it skips the cells that must be zero and the rows that can no longer reach the last one. Two samples of 49 five-point scores, which take the exact path by default, went from 8.6s to 0.026s (R 4.6.1 takes 10.2s), and 49 against 49 with a single tie from 10.3s to 1.15s. - The table is now built for the smaller sample, whose rank sum is the total less the other's: 45 against 5 went from 0.067s to under 1ms. - Against the previous build, 2400 random tied intervals and 900 exact p-values came out the same to the last bit. [wilcox_test(): the exact tables without ties] - Both are symmetric, and are now built and held as their low half only, which halves their memory and their time. The rank-sum table loops over the smaller sample: m = 5 against n = 3000 with exact => 1 went from 28ms to under 1ms. - The rank-sum recurrence is now bounded by the degree of the polynomial built so far, so it no longer forms counts as differences of numbers near C(m+n, n). Against exact rational arithmetic, two-sample p-values from large exact tables were off by a median of 15 ulp (the worst by 100780) and are now off by a median of 1 (the worst by 7). - The signed-rank table for ties or zeroes is symmetric as well, since flipping every sign turns V into sum(z) - V, and is folded the same way. The comment that said otherwise was wrong. [wilcox_test(): asymptotic intervals] - The root search re-ranked both samples from scratch at every step. Subtracting a shift and digits.rank's signif() both keep a sorted sample sorted, so x and y are now sorted once and each step merges them, which gives the same ranks in the same order: the interval on two samples of 20000 went from 52ms to 9ms. [wilcox_test(): infinite observations] - A difference of two infinities of the same sign (Inf - Inf) is NaN, and R's sort() drops it from the differences an exact interval and the Hodges-Lehmann estimate are read from. They were kept here, where no sort can place a NaN, and the estimate could come out Inf where R's is a number. - The scan behind an exact interval tries shifts that meet an infinite observation as Inf - Inf, and R's rank() ranks the NaN last. It was ranked wherever the sort left it: on SciPy's gh-11355 data wilcox_test gave [-5, 3] where R gives [-5, Inf], and 7 of 15 intervals differed. All 15 now match R, and t/wilcox_test.R.scipy.t pins them. - An infinite observation leaves the asymptotic interval no finite bracket. R's two-sample code dies in uniroot() with "invalid 'xmin' value", and its one-sample code returns a NaN interval at level 0. Both now return that NaN interval, with a warning, where the root search used to run on whatever the ranks of x - Inf came to. [wilcox_test(): tied arrays, and smaller fixes] - An element of a tied array has no value until its FETCH runs, which is get magic, and SvOK() was tested first, so a tied x croaked that it was empty. Each element is now fetched once. - Three croaks did not name the function; they now begin "wilcox_test:". - The exact signed-rank test with zeroes took its scale from every rank, the zeroes' included, where R takes it from the non-zero ones, so two tied zeroes doubled the table for nothing. - The exact intervals without ties sorted all m*n differences (or Walsh averages) for three or four order statistics, which are now placed by selection. - t/wilcox_test.R.scipy.t adds R's own tied exact-interval cases from src/library/stats/tests/ks-test.R, at full precision, and t/wilcox_test.t tied arrays and leak checks for every interval path. [write_table(): write errors, empty tables, ragged rows and tied data] - A write that fails now croaks and announces nothing. Output is buffered, so a full disk shows only at the close, which was never checked for a plain file, a LaTeX file or a workbook: a table written to /dev/full returned normally and said it had been written. R's own tests/reg-tests-1d.R (PR#17243) expects write.table() to complain there, and pandas raises. - The message gives the system's reason, as R's and pandas' do: "write_table: could not finish writing 'out.csv': No space left on device", and $! is left set to it. The same close meets an exceeded quota, an I/O error or a file server that has gone away, and only the system knows which it was. A compressed file's message, which was "could not write '...'" with no reason, is now the same, and a file that cannot be opened gives its reason too ("Permission denied", "No such file or directory"). - Empty data, {} or [], writes a file: the header col.names and row.names give, or a single empty record, as pandas writes ",A\n" for DataFrame({"A": []}).to_csv(). It used to return without writing or saying anything, which t/write_table.t and t/write_table.announce.t pinned; they now pin the new behaviour. - An array of arrays' row longer than its header lost its extra cells without a word. The header is now widened with empty cells, as pandas pads DataFrame([[1, 2], [4, 5, 6]]), and the long rows are warned about. - For a hash of arrays or an array of hashes, row.names => 'col' left the label column's header cell empty, so the name was lost and read_table read the column back as row_name. It is now headed 'col', as pandas heads a named index. - A hash of arrays took its row count from every array in the hash, so a col.names leaving out a longer one wrote rows of separators after the data ran out. It is now the longest of the arrays written. - A hash of arrays with a col.names naming no column croaked only after opening, and so emptying, the output file. It now croaks first. - A tied hash or array (Tie::IxHash, say), as a row, a column or the table itself, was written as empty cells, because SvOK() was tested before its FETCH ran. - A cell holding a NUL was cut short at it. Delimited output now writes it whole, raw and unquoted, as csv.writer and pandas' to_csv() do; .xlsx drops the NUL, which XML cannot hold. - The .xlsx sheet name is now checked as openpyxl checks one -- empty, or holding any of \ * ? : / [ ], is an error, and longer than 31 characters a warning -- and is written as UTF-8 when it arrives as Latin-1 bytes. A LaTeX table with no columns, which LaTeX rejects, is refused before the file is opened. - README.md said undef cells are written as NA by default. They are written empty, as t/write_table.t has always tested; the documentation now says so. [write_table(): memory and time] - .xlsx output is streamed: each row's XML goes to the file as it is made, and the worksheet's local header is patched once its size is known. The table used to be copied into SVs, then into one string of XML, then that into a second string holding the whole archive. A 200000 x 20 array of hashes took 1125 MB over the data's own peak for its 308 MB workbook and now takes 30 MB, and 2.31s became 1.52s. A pipe, which cannot seek, has its worksheet gathered in memory first. The bytes written are the same as before in every case, and a write that croaks partway now leaves a truncated workbook, as it always did a delimited file. - A number was formatted with SvPV(), which caches the text in the SV it is given: a hash of arrays holding a million numbers weighed 32 MB before a call and 96 MB after, and stayed that way. It is now formatted in a scratch copy. - Finding the columns of an array or hash of hashes made a mortal copy of every key of every row, all freed only when write_table returned: 202 MB for a 200000 x 20 array of hashes. They are now freed a row at a time, leaving 30 MB, the hash iterators that perl's own keys() leaves behind as well. - A hash of arrays looked each column's array up once per cell, and now does once per column: 0.209s became 0.155s on 200000 x 20. Each field is scanned once for what makes it need quotes, where it was scanned five times, and a quoted field is written in runs rather than a byte at a time. - Every buffer and handle is now on the save stack, so a croak from anywhere -- a tied FETCH included, which none of the dozen hand-written cleanups before each croak saw -- closes the file and leaks nothing. 0.3211 2026-09-27 CDT [min(), max(), sum() and the other reductions: magical arguments] - An argument whose value exists only once its get magic has run was read as undef. That covers `$#array`, a tied scalar whose FETCH had not yet been called, and a `substr()` lvalue: none has any value flags before `mg_get()`, and the reductions asked `SvOK()` first, so `sum(0, $#list)` croaked "undefined value at argument index 1" where List::Util's `sum()` returns the index. A tied scalar holding an array ref was likewise not seen as one. This affected `min`, `max`, `sum`, `mean`, `sd`, `var`, `median`, `mode`, `uniq`, `scale`, `skew` and `kurtosis`. Each now runs every argument's get magic once, on entry, and reads the fetched value from then on. - The same fault reached a tied element of an ordinary array in `median`, `skew`, `kurtosis`, `mode` and `scale`, which read the array's cells without running their magic. The first four croaked on it; `scale`, which skips undef, silently returned one value fewer than it was given. `sum`, `mean` and the rest already read such an element correctly. - A value is now fetched once. Reading one after its magic had run could run FETCH again, on every perl before 5.18 and in `mode` and `uniq` on all of them, and a tie that computes its value then answers differently the second time. `sd`, `var` and `scale` still read a tied array, or a tied element, once per pass, as they are documented to. - `median` on a tied array now croaks on a non-numeric element, as it already did on an ordinary one, instead of converting it to a number. `scale` croaks if a tied array answers differently between its counting pass and its value pass, which could previously write past the buffer the first pass sized. [Tests] - t/min.max.sum.ListUtil.t carries List::Util 1.70's own t/min.t, t/max.t and t/sum.t, from the copy bundled with perl 5.44.0. Their GETMAGIC block is what found the fault above. Where the two modules differ by design the header says so: `sum()` of nothing croaks rather than returning undef; a Math::BigInt is reduced as the NV its 0+ overload yields; and a sum is carried in an NV, not an IV. - An array reference is read as data whether or not it is blessed and whether or not its class overloads 0+, where List::Util would treat an overloaded object as a single number. This is a decision, not an oversight, so that an array-based object keeps being read as the values it holds, and the port now asserts it in place of upstream's three `example` cases. - t/get.magic.args.t checks all twelve functions with six kinds of magical input against the same call with plain values, and counts how often FETCH runs. [write_table(): .gz and .bz2 on Windows] - A compressed file is meant to hold exactly the text the plain file would, and on Windows it did not. The plain file is opened with perl's default layers, which there translate "\n" to CRLF; the :raw that makes way for the compressing layer removed that translation, so in 0.321 a .gz or .bz2 held LF lines where the same call to a plain file wrote CRLF. Where perl is a CRLF shop (PERLIO_USING_CRLF), the buffer above the compressor is now :crlf instead of :perlio, and the two files match again. Other platforms are unchanged. read_table reads either line end from a compressed file, as it does from a plain one. - t/write_table.compressed.pandas.t failed 16 of its 88 tests on an MSWin32 smoker for this. Its expected text now uses the line end a plain file actually gets, and the check of the "wrote" line reads its pipe in binary mode and ignores a CRLF from STDOUT's layers, which is perl's doing rather than write_table's. [lm()] - A non-integer exponent inside I() was truncated: the exponent was read with atoi(), so I(hp^0.5) became hp^0, a column of ones that was then aliased against the intercept and reported as NaN, with no sign that anything had gone wrong. I(x^1.5) became x^1. The exponent is now read as a number at the build's NV width, and one that is not a number, such as I(x^abc), is missing rather than 0. glm(), aov() and the other model functions read I() the same way and are fixed with it. - An infinite value in the data is refused with R's own message, "lm: NA/NaN/Inf in 'y'" (or 'x'), naming the row, as glm() already refuses it. It used to go into the fit, and every coefficient came back NaN, each reported with t = -Inf and p = 0: a test on the standard error being positive sent a NaN one to the infinite branch. NaN, which is missing, still drops its row. - t values, R^2, adjusted R^2 and F are the plain ratios summary.lm() takes, so 0/0 is NaN. A response with no variation at all reported t = Inf, p = 0 and R^2 = 0 for every coefficient; it now reports NaN, as R does, and a nonzero estimate over a standard error of exactly 0 is +-Inf with p = 0. - The residual degrees of freedom are now tested against the rank, not the column count. y ~ x + z on three rows with z = 2x has one residual degree of freedom, and R fits it with z NA; lm() refused it as "0 degrees of freedom". A fit with none at all is still refused. - Everything lm() allocates is on the save stack, so the new croaks free it all. - t/lm.edge.R.t covers it, from R 4.6.1 on its own mtcars and from SciPy 1.18.0's test_regressZEROX (Wilkinson's W.IV.D); t/lm.edge.R.R regenerates the reference values. Where R's QR leaves rounding in an exactly zero residual sum of squares and so reports t = 9.0e15 instead of Inf, the test pins the definition and says why. [Row names that are not ASCII] - fitted.values, residuals and the other hashes keyed by row name, from lm(), glm(), zerotrunc(), hurdle(), svyglm(), ivreg(), lmer(), aov() and predict(), dropped the UTF-8 flag of a name, so "\x{65e5}\x{672c}" came back as its six bytes and the caller's own key did not find it. A HoH key that fits in Latin-1 happened to survive, but the same name given in row.names or _row did not. Row names are now held as UTF-8 and stored as UTF-8 keys, which perl downgrades where it can, so ASCII and Latin-1 byte names come back exactly as they were given. - t/rownames.utf8.t checks every one of those functions with each data shape that carries names. [min(), max(), sum(), mean(), sd() and var(): faster on plain arrays] - The loops that walk an ordinary array of numbers now keep four running values instead of one, and combine them at the end. With one, every add or compare waited on the result for the previous element, and for min() and max() that wait was most of the cost: they ran at 1.06 ns/element on 1e4 values where sum() ran at 0.67, and a plain C loop over a flat array of doubles was no faster. sd() and var() share the first pass of sum(), so they gain too, by less. - Times are nanoseconds per element, 0.321 -> 0.3211, on perl 5.44.0 at -O2, one pinned core, best of 120 timings taken in alternating runs of the two releases, on the same normal data. At n = 1e6 the data no longer fits in cache, so that column is memory-bound and its cells moved by up to 10% between runs; the others moved by under 1%. - | function | n = 1e3 | n = 1e4 | n = 1e5 | n = 1e6 | |---|---|---|---|---| | `max` | 1.094 -> 0.905 (-17%) | 1.063 -> 0.855 (-20%) | 1.060 -> 0.850 (-20%) | 1.461 -> 1.160 (-21%) | | `min` | 1.094 -> 0.906 (-17%) | 1.063 -> 0.854 (-20%) | 1.064 -> 0.850 (-20%) | 1.392 -> 1.176 (-16%) | | `sum` | 0.712 -> 0.617 (-13%) | 0.670 -> 0.572 (-15%) | 0.721 -> 0.578 (-20%) | 1.295 -> 1.064 (-18%) | | `mean` | 0.712 -> 0.617 (-13%) | 0.670 -> 0.571 (-15%) | 0.721 -> 0.583 (-19%) | 1.325 -> 1.048 (-21%) | | `sd` | 1.436 -> 1.352 (-6%) | 1.384 -> 1.286 (-7%) | 1.488 -> 1.332 (-10%) | 2.700 -> 2.596 (-4%) | | `var` | 1.436 -> 1.378 (-4%) | 1.383 -> 1.293 (-6%) | 1.483 -> 1.357 (-9%) | 2.739 -> 2.542 (-7%) | - sum(), mean(), sd() and var() now add the elements in a different order, four interleaved partial sums added pairwise, so a result can differ from 0.321's in the last bits. The worst-case rounding error is smaller, since each partial sum is a quarter as long. - min() and max() now order -0 below +0, as IEEE 754-2019's minimum() and maximum() do: max(-0.0, 0.0) and max(0.0, -0.0) are both +0, and both min()s are -0. They used to return the first zero they met, as R does, and with four lanes that would have come to depend on which lane each zero fell in. NumPy and List::Util return the last zero they meet, so none of the three agrees with this, and t/min.max.signed_zero.t records all of them. The perl literal -0 is the integer 0, which has no sign, so max(-0, 0) was always 0. The table above includes the cost of tracking the sign, which is 0-6% on normal and integer data and up to 9% on data that is mostly zeros. - An array holding a numeric string, a tied element or a hole is still read exactly as before: the fast loop stops at that element and the ordinary one finishes from there. 0.321 2026-09-25 CDT [read_table(): warnings about stray quotes] - A '"' that opens a quoted field takes every newline into that field up to the next '"'. <=0.320 nothing said when that had happened: a file with one unclosed '"' came back with the rest of the file in a single cell, and a '"' used as text in the middle of a field (an inch mark, 5'10") ran rows together or ended in an "Alignment error" that said nothing about quotes. read_table cannot tell a stray '"' from CSV quoting by the bytes alone, and neither can R or pandas, so it now says what it saw and names the line the quote opened on. - A file that ends inside a quoted field is warned about, and what was read is kept, as R's scan() does ("EOF within quoted string"); pandas raises instead. That cell no longer gains a newline the file did not have when the file has no final one. - A '"' in the middle of a field that opens a quoted field running past the end of its line is warned about, once per file. It is still read as a quote, as R reads it (tests/reg-tests-1d.R pins '="Total' opening one); pandas keeps such a '"' as text. A quoted cell that starts at the start of its field and holds a line break is ordinary CSV and is not warned about. - An alignment error on a row that a quoted field ran across lines in now says where that quote opened and whether it was ever closed, on the XS fast path and through a filter alike. - quote => '' warns when every field on the first line is wrapped in '"', as R's write.csv() writes them, since the quote marks would then stay in every name. An .xlsx is not checked. - t/read_table.quote.R.pandas.t covers all of it, from pandas' test_quoting.py, test_skiprows.py, test_eof_states and GH 62739 and R 4.6.1's reading of the same inputs; t/read_table.quote.R and t/read_table.quote.pandas.py regenerate the reference values. [write_table(): .gz and .bz2 output] - A file name ending in .gz or .bz2 is written gzip- or bzip2-compressed, holding exactly the text the plain file would. The rest of the name picks the default sep as before, so x.tsv.gz is tab-separated. It is streamed through a PerlIO::via layer at gzip's and bzip2's default levels, and a gzip header carries no name or time, so the same table always makes the same bytes. Up to 0.320 such a name got plain text. - A write that croaks partway leaves a truncated file, which read_table refuses, never a whole-looking one; for bzip2 the first block is written at once so that even a croak on the first row leaves one. A compressed write that cannot reach the disk croaks; a plain write's errors are still not checked. - Only delimited text is compressed: .tex.gz, .xlsx.bz2, or tex/xlsx with a compressed name, is an error. So is a .bgz name, which promises bgzip's BGZF, which this does not write. - t/write_table.compressed.pandas.t covers it, from pandas' test_compression.py and test_to_csv.py. [Minimum perl] - perl 5.10.1 is now the oldest supported, and Compress::Raw::Bzip2, core from that release, is a declared prerequisite. [read_table(): gzip and bzip2 input] - A gzip or bzip2 file is read as the text inside it, with nothing to ask for. It is recognised by its first bytes, not its name, as R's file() recognises one for read.table, so a compressed file without a .gz suffix is read, and a plain one that has one is still text. The extension still picks the default sep, from the name inside: x.tsv.gz is tab-separated. Up to 0.320 a compressed file was parsed as if its bytes were text, and came back as garbage rows or an alignment error. - The file is inflated as it is read, 64 KB at a time, through a PerlIO::via layer over Compress::Raw::Zlib or Compress::Raw::Bzip2, so a large .tsv.gz takes no more memory than the plain file would; the gzip read of a 110 MB CSV costs about what zcat does on top of reading it plain. - Every member is read: bgzip's BGZF (every .vcf.gz), R's gzfile(, "a") and `cat a.gz b.gz' all write several, and stopping at the first would have dropped all but the first block without a word. A truncated file, a bad CRC, and data after the last member that is not NUL padding are errors that name the file, never a short read. - Both need only core modules. Compress::Raw::Bzip2 is loaded only when a bzip2 file is read, and a read without it says what is missing. - t/read_table.compressed.R.t covers it, from R 4.6.1's tests/reg-tests-1b.R compressed read.table and append-mode cases and pandas' test_compression.py; t/read_table.compressed.R writes the fixtures. [read_table(): rows of empty fields, and two header fixes] - A line holding nothing but separators is a row of empty fields, as R 4.6.1's read.table and pandas 3.0.4's read_csv both read it. Up to 0.320 a tab-separated "\t" was taken for a blank line and skipped, so a TSV row with every cell empty went missing without a word and the row count no longer matched R's or pandas'. A line of blanks none of which is the separator is still skipped, and so is any line of blanks under qr/\s+/. t/read_table.blank_lines.R.pandas.t covers it, from pandas' test_empty_lines and test_whitespace_lines; t/read_table.blank_lines.R and t/read_table.blank_lines.pandas.py regenerate the reference values. - auto.row.names with a commented-out header ("# a\tb") refused the header, because the data rows are one field wider than it -- the very shape auto.row.names looks for -- and made the first data row the header instead. The header is now kept and the rows named. - A worksheet name in an .xlsx is decoded by the one-pass decoder the cells use. It went through five substitutions in turn before, so a reference one of them produced was decoded again by a later one ("&lt;" became "<"), and a numeric reference too large for chr() died. [read_table(): speed] - Once the XS fast path is building the rows, an empty field is parsed straight to the undef it becomes, instead of to an empty string that was then freed and replaced. On a 300,000 x 10 CSV with nine cells in ten empty, an aoh read goes from 0.185 s to 0.106 s, a hoa from 0.164 s to 0.097 s and an aoa from 0.152 s to 0.073 s; a file with no empty cells is unchanged. A filter is still handed '' for an empty field, as before, and an .xlsx gap is treated the same way. - na.strings is checked by comparing the few strings given rather than by hashing every cell: with three of them, an aoa read of a 300,000 x 10 CSV goes from 0.155 s to 0.134 s. Past eight strings the hash is used as before. [Tests] - t/read_table.aoa.t failed 3 subtests on Windows (Strawberry Perl 5.42.2): write_table writes its file in text mode, so its lines end "\r\n" there, as R's write.csv() does, and the test read the result back raw. It now reads it through the platform's text layer. Nothing in the module changed. 0.320 2026-09-25 CDT [read_table(): sep may be a qr// regex] - sep (and its synonym delim) was only ever a literal string. A qr// was stringified to "(?^:\s+)", matched byte for byte, never found, and every line came back as a single field keyed by the whole header line, with no warning. A qr// is now a pattern: sep => qr/\s+/ reads whitespace-aligned columns, qr/\s*,\s*/ trims around each comma, and qr/[;,]/ takes either character. A string sep is the literal it always was, so sep => '\s+' still means those three characters. - qr/\s+/ is whitespace-delimited, as sep=r"\s+" is in pandas and sep = "" in R's read.table: leading and trailing whitespace on a line make no field. Any other pattern cuts as split() does, so a separator at the start of a line leaves an empty first field. Capture groups in the pattern are not fields, and a pattern that can match an empty string, such as qr/\s*/, is refused rather than cutting between every character. The pattern is matched as written, so its group numbers are its own and a backreference works: qr/(:)\1/ splits on "::". - Quoted fields, comments and commented-out headers, blank lines, a byte-order mark, CRLF, filter, row.names, auto.row.names, na.strings and all three output types behave exactly as with a literal separator: the same XS parser reads both, and only how it finds a separator differs. pandas drops quote handling for a regex separator; read_table keeps it. An .xlsx still ignores sep. - The separators are found by perl's regex engine, called from C on the parser's own line buffer, so a regex read keeps the C fast path: 0.17 s rather than 0.14 s for a literal sep on a 300,000 x 5 CSV. A literal sep is untouched and as fast as before. [read_table(): header => 0, col.names and quote => ''] - read_table always took the first line as the header, so a file without one -- an NCBI taxonomy dump, R's write.table(col.names = FALSE) -- lost its first row to the column names. header => 0 (R's header = FALSE, pandas' header=None), or perl's false '', reads it as data, and the columns are named by col.names or, as R names them, V1, V2, ... col.names with a header renames its columns, warning as R does when the lengths differ. With no header there is none to rescue, so a line starting with the comment marker is a comment even when text hugs it ("#comment"), as in R. - A '"' anywhere in an unquoted field opened a quoted one, so a file whose quotes are not CSV quoting had every line up to the next '"' read into one cell, silently. NCBI's fullnamelineage.dmp has a name ending in 'Beach rock 4+5"', and about 950,000 of its 3,015,956 lines went into one field. quote => '' (R's quote = "", pandas' quoting=QUOTE_NONE) makes '"' ordinary text, with a literal sep and a regex one alike; that file now reads in full. The default is unchanged. [write_table(): a hash of hashes keeps its outer keys] - A HoH's outer keys were written only as row names, and row.names has been off by default since it stopped following R's write.table. For every other shape that label is a 1..n index nobody needs, but a HoH's keys are the only place its row identifiers exist, so a taxid-keyed HoH came out with no taxids in it at all. A HoH now defaults row.names on, under an empty header cell, which read_table reads back as row_name. row.names => 0 still drops the keys, and every other shape is unchanged. - row.names => 'name' on a HoH was accepted and did nothing a 1 would not: the key column's header stayed empty. It now heads that column, so row.names => 'taxid' writes "taxid,genus,species", in delimited, LaTeX and .xlsx output alike. A name that is also a column being written -- a col.names entry, or else a key of any inner hash -- dies before the file is opened, rather than writing two columns of one name. [read_table(): output.type => 'aoa'] - read_table could return an AoH, a HoA or a HoH, but not the AoA that write_table, agg, csort and melt all take. output.type => 'aoa' returns the header row, then one array per data row, every row in file column order, with an empty or na.strings cell undef as in every other shape. The header row is what write_table reads an AoA's first row as, so a read and a write round-trip a delimited file byte for byte. - It is the one shape that keeps every field when the header repeats a name, so it gives no "later values win" warning. Nothing labels an AoA's rows, so row.names with it is an error. It goes through the same C fast path as the other three shapes when there is no filter, and through the per-row closure when there is one. [Tests] - t/read_table.regex_sep.t takes its whitespace and multi-character cases from pandas 3.0.4's parser tests, and checks that every CSV under t/, and a tab-separated copy of each, reads the same through qr/\Q$sep\E/ as through the literal separator, warnings and errors included. - t/read_table.header_quote.R.pandas.t takes its cases from R 4.6.1's tests/reg-IO2.R and reg-tests-1a, 1b, 1d and 2, with the values R itself gives frozen from t/read_table.header_quote.R, and from pandas 3.0.4's test_header.py, test_quoting.py, test_dialect.py and test_na_values.py. The one cell where R and read_table differ -- R ends a line at a comment character anywhere in it -- is pinned there. 0.319 2026-09-22 CDT [glm(): offsets, prior weights, robust covariance, absorbed factors] - glm() took only formula, data, family, theta and conf.level, and died on anything else, which left count models unable to put a rate on person-time. It now takes offset() terms in the formula and an offset argument (a column, an expression such as 'log(t)', or an array ref), and the negative-binomial theta search and the null deviance both see the offset, as MASS::glm.nb()'s do. - weights are R's prior weights. A binomial fit whose weights make a non-integer number of successes warns as R does. - vcov => 'HC0' .. 'HC3' gives sandwich::vcovHC()'s covariance and cluster gives vcovCL()'s, one to four ways ('firm + year'), with every standard error, z, p-value and interval recomputed from it. A poisson fit on a 0/1 outcome with HC0 is the modified-Poisson risk ratio. - A factor after '|' in the formula, or in absorb, is absorbed by weighted within-group demeaning instead of expanded into dummy columns, so a factor with thousands of levels costs a vector per level rather than a column. Groups whose outcome sits at a boundary are dropped, as fixest::feglm() drops them, and fe.removed counts their rows; only rows with a positive weight decide that, so a zero-weight count cannot keep an all-zero group in the fit. - The IRLS loop was rewritten around one core shared with the new models below, with maxit and epsilon exposed and R's convergence rule and penultimate-iterate standard errors. New results: loglik, vcov, vcov.type, dispersion, nobs and, where they apply, n.clusters, absorb, fe.removed, twologlik and SE.theta. predict() re-evaluates an offset on new rows, and croaks on a model it cannot predict from (an offset given as an array, or absorbed factors) rather than leaving the term out. - The formula reader now evaluates log(), exp(), sqrt(), log2(), log10(), log1p() and abs() of a column, which offsets and ivreg's examples need. - Validated in t/glm_offset_weights.R.t, t/glm_vcov.R.t and t/glm_absorb.R.t against R's glm, MASS, sandwich and fixest test suites, statsmodels' and Stata's pinned results, and a full-dummy fit. [glm(): Inf, a NaN in step-halving, and large negative-binomial counts] - An infinite response, covariate, offset or weight is not missing, so its row went into the fit and every coefficient came back NaN without a word. glm() now croaks "NA/NaN/Inf in 'y'" (or 'x', 'offset', 'weights') with the row name, as R's Cdqrls() stops. - A coefficient aliased on one IRLS iteration was carried into the next as NaN, so if the column stopped being aliased and the step was halved, the NaN reached the linear predictor. It is carried as 0, as glm.fit() does. - The negative-binomial log-likelihood summed y logs per row for every count below 1e6, so a fit cost O(sum y): 20000 counts near 34000 took 11.7 s against 0.01 s for the Poisson fit. Counts from 64 up now go through Loader's stirlerr(), and the fit takes 0.03 s. It is also more accurate: 3e-13 relative against mpmath at 60 digits, where the long sum was 1.9e-9 out. t/glm.t pins it. - X'WX is accumulated over its upper triangle only and skips zero-weight rows, halving the IRLS loop's main cost and leaving the matrix exactly symmetric; coefficients move by about 1e-12. Absorbed factors are demeaned all columns in one pass over the rows, bit for bit as before, and vcovHC() no longer holds an n x p copy of the scores. [coxph(): counting-process data, strata, robust variance, a formula] - coxph() took only (\@time, \@status, covariates), which cannot express a time-varying covariate or late entry. It now takes (start, stop] data, strata, a cluster with the grouped-dfbeta robust variance, case weights and an offset, either as options to the positional form or through formula => 'Surv(start, stop, event) ~ x + strata(g) + cluster(id)' over a data set. Intervals that span no event are skipped, as survival's agreg.fit skips them. New results: var and vcov, the score and Wald tests, and naive.se, naive.var, robust.score.test and n.clusters under a robust variance. - Validated in t/coxph_extended.R.t against survival's own tests (bladder, the phreg corpora) and statsmodels'. [New models: zerotrunc, hurdle, svyglm, ivreg, lmer] - zerotrunc() is countreg::zerotrunc(), a Poisson, negative binomial or geometric regression truncated at zero, fitted by damped Newton on the exact likelihood and its analytic Hessian. hurdle() is pscl::hurdle(), with a logit or a censored count zero part and the regressors after '|' for it. t/zerotrunc_hurdle.R.t pins both to countreg and pscl, with an mpmath third opinion at 60 digits (t/zerotrunc_hurdle.mpmath.py). - svyglm() is survey::svyglm() on a one-stage design: sampling weights, strata, PSUs and a finite-population correction, with the Taylor linearisation variance and design degrees of freedom. t/svyglm.R.t is taken from survey's tests on the api data. - ivreg() is ivreg::ivreg(): two-stage least squares by Householder QR, two- or three-part formulas, weights, HC0/HC1 and clustering, and the weak-instrument, Wu-Hausman and Sargan diagnostics. t/ivreg.R.t follows ivreg's tests and examples and statsmodels' Stata ivreg2 results. - lmer() is lme4::lmer() by REML or ML, with random intercepts and slopes, correlated or not, crossed or nested grouping factors, and lmerTest's Satterthwaite degrees of freedom. The fit is held to a tightly converged lme4 fit rather than to lme4's default one, which stops about 1e-6 short in theta. t/lmer.R.t covers lme4's and lmerTest's examples and statsmodels' mixed-model corpora. [anova() compares fitted models] - anova($m0, $m1, ...) with lm or glm fits is R's anova.lmlist / anova.glmlist, with test => 'F', 'Chisq' or 'LRT' and dispersion, and with negbin fits MASS's anova.negbin likelihood-ratio table; data and formulas still go to the XS anova() as before. The dispatching wrapper has no prototype, because the XS one's ($@) would put anova(@fits) in scalar context. Validated in t/anova_fits.R.t against R and MASS. [read_table: five bugs] - The CSV parser split lines on $/, not on newlines, because sv_gets() reads PL_rs. Under a `local $/;` anywhere up the call stack the whole file was one "line" and read_table returned [] without a word; a record length ($/ = \N) cut rows at N bytes. The parser, and the perl-side peek that recovers a commented-out header, now split on "\n" whatever $/ is, and a filter still sees the caller's $/. t/read_table.input_record_separator.t. - A UTF-8 byte-order mark, which Excel's "CSV UTF-8" export writes, was read as part of the first header name ("\xEF\xBB\xBFid"), and hid a comment or commented-out header behind it. It is dropped from the start of the file, as pandas drops it; t/read_table.bom.pandas.t takes its cases from pandas' test_utf8_bom and test_first_row_bom. - In an .xlsx, a cell with formatting and no value () past a row's last value widened every row to reach it, so a shaded column came back as unnamed columns of undef and a duplicate-name warning. Such cells no longer count toward the width, which is what readxl 1.5.0 and pandas 2.2.3 give for the same workbook. - A hoh read of a file whose only column is the row name came back as {}: the per-row hash was only made when a value was stored in it. Each row is now an empty hash, as R's read.table gives n rows of 0 columns. A read error part-way through a CSV is now an error rather than a truncated table. [read_table: faster] - 'hoh' now goes through the XS fast path that aoh and hoa already took, with the perl closure's duplicate-row-name warning and undefined-row-name error spelled the same; a 300,000 x 5 CSV reads in 0.20 s rather than 0.91 s. Only a 'filter' still needs the closure. - An unquoted field is copied straight from the line buffer instead of through the field accumulator, and the row hashes are filled through shared-hash-key copies of the column names, so no key is hashed again per cell. The same file reads as an aoh in 0.096 s rather than 0.120 s. [write_table: wide characters in row.names and tex.longtable.head] - write_table() croaked "Wide character in subroutine entry" when row.names named a column outside Latin-1, or when tex.longtable.head was such a caption: the XS check for a non-numeric value read it with SvPVbyte. It now reads the bytes as stored, and no longer downgrades the caller's string in place or hands isdigit() a negative char. Found by XS::Check; t/write_table.t and t/write_table.longtable.t pin both. [dunn_test, p_adjust: method names fold by ASCII rules] - The method name was lowercased with tolower(), which is undefined for a byte >= 0x80 in a signed char and follows LC_CTYPE, so under a Turkish single-byte locale 'BONFERRONI' folded its I outside ASCII and was rejected. Both now fold A-Z only. t/dunn_test.t and t/p_adjust.R.t. 0.318 2026-09-18 CDT [transpose() took the interpreter down on a tied array of arrays] - The array branch walked its input twice: one pass validated every row as an array ref of row 0's width, then one pass per output column re-fetched each row to copy a cell out of it. The copying pass tested what the second fetch returned -- `if (elem && *elem) SvGETMAGIC(*elem);` -- and then dereferenced it regardless, `(AV *)SvRV(*elem)`, without re-testing `SvROK`. - A tied array's elements do not exist until `FETCH` has run, so the second read of a row is a fresh call that may hand back something other than what the first one did. `SvRV()` on a plain string then takes the PV's buffer for an `AV *` and the walk runs off into it: SIGSEGV, which `eval` does not see. A two-row tied array whose `FETCH` stops returning an array ref once the validating pass is over reproduces it in one call. Nothing a plain array could hold reached it, because nothing between the two passes could change a row. - There is one pass now, so the read that is validated is the read that is copied, and the check that rejects `transpose([1, 2, 3])` rejects this too. [transpose() on an array-of-arrays: the loop nest was column-outer] - Every cell cost two out-of-line `av_fetch()` calls -- one to re-reach row i through the outer array, one for the cell itself -- plus an `av_store()` and a second round of get magic. A 300,000 x 3 frame fetched the outer array 900,000 times for the 300,000 rows it has, and a 950 x 950 one walked the whole outer array 950 times over. It is the same shape `cor_extract_cols()` and the AoH-to-HoA reader in this file already record having turned inside out. - Row-outer costs one fetch of each row. The output columns are allocated at their final length, so nothing is grown or copied, and their blocks are zeroed with `AvFILLp` set to the last row up front rather than advanced behind each store: one store per cell instead of two, and still safe to croak out of part-way through, because a slot not yet written is a `NULL` that `SvREFCNT_dec()` ignores. The `Zero()` is what makes that true on 5.10, where `av_extend()` fills the slots it adds with `&PL_sv_undef` rather than the `NULL` it has used since 5.20. - On perl 5.44.0 at -O2, best of seven, at 900,000 cells throughout: 300,000 x 3 went 15.8 ms to 5.8 ms, 3,000 x 300 went 17.9 ms to 4.5 ms, and 950 x 950 went 14.4 ms to 3.7 ms. [transpose() on a hash-of-hashes: a mortal key SV per cell] - The cell loop called `hv_iterkeysv()` for the column key once per cell, and a mortal is not reclaimed until the XSUB returns. A 3,000 x 100 frame held 300,000 of them for the length of the call: 29.1 MB of resident memory to produce a result of about 14.5 MB. Both keys are read straight out of the HE now -- `HePV()` for the bytes, `HeUTF8()` for the flag, `HeHASH()` for the hash -- so a transpose allocates no key SV at all, and the same frame peaks at 14.6 MB. - Handing perl a precomputed hash together with a negative (utf8) klen is safe because perl invalidates the hash itself when it has to downgrade such a key to bytes, in hv.c's `HVhek_KEYCANONICAL` block; `row_drop()` already relied on that. A tied hash iterates with SV keys, which have no HEK and so no `HeHASH()`; those pass 0 and let perl compute one. - Each output column is also pre-sized with `hv_ksplit()` to the number of input rows it will end up holding, rather than splitting its way there and rehashing everything already in it once per doubling. That one call is most of this branch's time: the same 3,000 x 100 frame takes 33.8 ms without it and 20.2 ms with, against 30.7 ms in 0.317. [transpose(): what did not change, and two smaller fixes] - Nothing about what it returns moved. Cells are still shared by refcount rather than copied, a ragged array still croaks, a physical hole still reads as undef, a tied input still reads through its magic, and the croak messages are the same text. - Those messages formatted a `size_t` with `%d` and a cast to `int`, which truncates above 2**31; they use `%" UVuf "` now. `av_len()` gave way to `AvFILL()`, which is the same answer with the untied case inline. [t/transpose.t] - Nine new subtests, 43 to 52. A well-behaved tied outer array and a tied inner row, both of which 0.317 handled and which are here so that the fast path cannot quietly stop handling them. The tied row that stops being an array ref, which is the SIGSEGV above, and which fails 0.317 as a signal rather than as a test. A tied hash, whose keys iterate as SVs and so take the `HEf_SVKEY` path through the new key reads. utf8 row and column keys, including one that perl downgrades to bytes on the way in, which is the case that says the hash handover is right. And a 20-wide frame whose 51st row is one cell short, which croaks with most of the result already written. Two of the nine are leak checks, on that part-filled croak and on the utf8 keys. - All eight perls in the local matrix pass, warning-free: 5.10.1, 5.12.5 (long double), 5.16.3 threaded long double, 5.42.3 threaded, 5.44.0, 5.44.0 with x87 arithmetic, 5.44.0 32-bit (`ivsize=4`), and 5.44.0 quadmath. 0.317 2026-09-12 CDT [srand() reached nothing this module drew, on a threaded perl before 5.20] - A CPAN smoker running perl 5.18.3 (`x86_64-linux-thread-multi-ld`: `useithreads`, `long double` NV) failed 0.316 in `t/rbinom.dist.t`, two subtests out of ninety. Both were reproducibility checks -- the same seed drawn twice, and the `prob`/`1-prob` mirror -- and every distribution test in the file passed, there and everywhere else. - On a perl older than 5.20, `Drand01()` is libc's `drand48()` and `seedDrand01()` is `srand48()`. When that perl is threaded, `reentr.h` redefines both over the interpreter's own `PL_reentrant_buffer` so that two threads do not share libc's global state -- but the whole block is wrapped in `#if PERL_REENTR_API == 1`, which `reentr.h` sets for `PERL_CORE` and `PERL_EXT` and nothing else. `perlxs` says as much under "Thread-aware system interfaces", and warns there that mixing the `_r` and `_r`-less forms of one interface is not well defined. - An XS file outside the core is neither, so `pp_srand()` seeded `PL_reentrant_buffer->_drand48_struct` while every draw in `LikeR.xs` read libc's process-global state, which nothing ever called `srand48()` on. `srand($seed)` had no effect whatever on `rbinom()`, `rnorm()`, `runif()` or `sample()`: two identical calls returned consecutive stretches of one stream that was never reset, and the first number a process drew came off an uninitialised `struct drand48_data` (3.9e-14, measured). The distribution was never wrong, which is why two reproducibility subtests were the only thing that ever failed. - `LikeR.xs` now carries `reentr.h`'s own two macros for `drand48()` and `srand48()`, copied verbatim and under the same prototype guards, so `Drand01()` and `seedDrand01()` expand to the state perl seeds. Perl 5.20 replaced the libc call with `Perl_drand48()` and `reentr.h` stopped wrapping `drand48` in the same release, so `PL_random_state` existing is the test for a perl that was never affected; on an unthreaded perl `USE_REENTRANT_API` is undefined and none of it applies. `random()` is left alone deliberately: `reentr.h` reads `_random_struct` and seeds `_srandom_struct` there, two different buffers, so on a perl configured that way `srand()` does not reach perl's own `rand()` either. [t/srand.stream.t] - New, and the reason the bug was module-wide but showed up as two subtests of one file: `t/rbinom.dist.t` was the only test in the suite that ever drew one seed twice. Fifteen checks that a seed determines what `runif()`, `rnorm()`, `rbinom()` and `sample()` return, that a different seed moves them, and that a draw made in XS and a draw made by perl's `rand()` come off one stream in order. Nine of the fifteen fail against 0.316 on an affected perl. No value is pinned: which numbers a seed gives is libc's business on some builds and perl's on others. [t/arg.crash.regressions.t ran its child perls through a sh command line] - A Strawberry perl 5.42 smoker (`MSWin32-x64-multi-thread`) failed 0.316 with 26 of that file's 98 subtests, which is every subtest in it that spawns a child. `child_result()` built the child's command as `` `$^X @inc -e '$prog' 2>&1` ``, and `cmd.exe` does not treat `'` as a quote character: perl was handed `'use` as its `-e` program and answered `Can't find string terminator "'" anywhere before EOF at -e line 1.` The snippets that carry a `"` failed differently and silently, as `exit 255 ()`. Nothing was wrong with the module -- every case those 26 subtests guard was still being rejected correctly. - The snippet now goes to the child in a file and the child is spawned with the list form of `system()`, so there is no shell and nothing in the snippet has to survive one. That matters beyond the quote character: the snippets hold `'`, `"` and a literal NUL between them, and no single quoting of a command line is right for both `sh` and `cmd.exe`. The child's verdict comes back in a file of its own whose path is passed in `@ARGV`, so it does not depend on inheriting a redirected handle either; the child's output is still captured, but only for the diagnostic. - It was the only test in the suite that shelled out, and the whole of what the smoker reported. [A threaded perl in the local matrix] - There was no threaded perl *older than 5.20*, which is what this needs, so no run here could have caught it: `perl-5.42.3` is threaded but is a 5.20-or-later perl, where `Perl_drand48()` has replaced the libc call and `reentr.h` no longer wraps `drand48` at all. `perl-5.16.3` built with `-Duseithreads -Duselongdouble` -- `USE_ITHREADS`, `USE_REENTRANT_API`, `randfunc=drand48`, the same three that matter on the smoker -- reproduces its two failures exactly, and is what the fix was tested against. `./test.all.perls.pl` picks it up like any other perlbrew perl. 0.316 2026-09-11 CDT [read_table on .xlsx: the worksheet parser moved from perl to XS] - Reading an .xlsx was slow out of all proportion to the format. On a 21,845 x 50 workbook (7.8 MB on disk, 36 MB of worksheet XML, 1,013,220 cells, an 8 MB shared-string table of 117,870 entries) `read_table` took 2.47 s and peaked at 263 MB. The same table written out as a 34.7 MB CSV and read back through `_parse_csv_file` took 0.139 s and 156 MB -- from a larger file, into the same array of hashes. Almost the whole gap was the perl side, not the format. - Where it went: the nested regexes in `_parse_xlsx_sheet` were 1.68 s of the 2.47 s, or 1.66 us per cell, against 0.099 s for a bare `$ws =~ m{ 1<<20` is not the way to do this: 2.08 s.) - `xl/sharedStrings.xml` is parsed in XS too, and `_xlsx_col_idx` is now the `xlsx_ref_col()` that places the cells, exposed to perl so that `t/xlsx_col_idx.t` still exercises the letter arithmetic on its own. - Together: 2.47 s -> 0.48 s and 263 MB -> 221 MB on that workbook, and 0.083 s -> 0.031 s on a 2,207-row one that uses inline strings and no shared-string table. What is left of the peak is the table the caller asked for, plus the 36 MB worksheet part and the 33 MB shared-string table, which are both live for the length of the parse. [What the new parser answers differently] - Two things, both on input the format does not allow. A cell reference past `XFD` -- the last of the 16,384 columns ECMA-376 gives a worksheet -- is now read as no reference at all, which puts the cell in the next column, the same place a cell with no `r=` goes. Every row is padded to the widest column the sheet mentions, so the reference is what decides what a row costs in memory, and a bad one should cost the 16,384 the format allows rather than the twelve million `ZZZZZ` asks for or the three hundred million of `ZZZZZZ`. The perl parser placed the cell wherever the arithmetic landed. - A numeric character reference above `�` is now left in the text instead of being decoded. The perl version handed the number to `chr()` unguarded, and that was not one answer but two: `�` came back as thirteen bytes of perl's extended UTF-8 on an ivsize=8 build, and died outright on `5.44.0-i686` with "Use of code point 0xFFFFFFFF is not allowed". XML 1.0 does not allow a character reference above `#x10FFFF` at all, so nothing legal is lost, and the answer is now the same on every perl in the matrix. - Everything else is unchanged, and was checked rather than assumed: the two parsers were run side by side over `Affinity Dataset(main).xlsx` and `titanic.xlsx` in all three `output.type` shapes with and without a `filter` and `na.strings`, over fifteen hand-built worksheets covering the awkward layouts, and over 400 generated strings of entity soup. Every `Data::Dumper` of the result is byte-identical, warnings and die messages included. [How the new parser was checked] - A tokenizer is only as good as the malformed input it survives, so it was fuzzed rather than reasoned about: every one of the 362 prefixes of a worksheet part, 20,000 byte-level mutations of it, and 12,000 generated workbooks each read twice -- once through the XS fast path and once through the perl callback -- so that the two could be required to agree cell for cell. They do, on all of them. Three defects in the new code turned up that way and are fixed: - `av_store()` already releases whatever was in the slot, so the parser's own `SvREFCNT_dec` for a repeated `r=` in one row was a second free of the same SV. Only malformed input has two cells claiming one column, which is why nothing but a fuzz would have found it. - The two passes could see different cells. The second skipped each cell body to ``; the first had no reason to and did not. On a cell whose `` is missing that search runs on to the next cell's, so the pass that skipped saw fewer cells, and rows were then built to a width measured from a different reading of the same file. Both passes now skip -- the scan is the same scan or it is not the same file. - The alignment croak that followed segfaulted. `S_fast_row()` frees the row buffer itself before it croaks, because it unwinds past `_parse_csv_file`'s local, and the new parser's save-stack destructor then freed it again. It now hands the buffer over and takes it back only if the call returns. - A fourth was found by the new test on `perl-5.10.1` and `perl-5.12.5`, and by nothing else: a cell is placed by its column reference, so a row with a gap leaves the slots between untouched, and perl 5.10 and 5.12 fill a newly extended array with `&PL_sv_undef` where 5.14 and later zero it. `SvCUR()` on the immortal undef dereferences a NULL `SvANY`. Only the first row after a fresh buffer can meet one -- every later row is cleared to NULLs -- and `t/read_table.xlsx.t`'s sparse cells are not in its first data row, so it passed on both perls while the new file segfaulted. The hole test now accepts either value. [read_table on .xlsx: a third off the time, straight to Compress::Raw::Zlib] - With the worksheet parser in XS, decompression is what a read of an .xlsx spends its time on: 0.239 s of the 0.413 s the 21,845 x 50 workbook takes, against 0.015 s to parse its 117,870 shared strings. `_unzip_member_fast()` reads the archive's central directory itself and inflates one member with `Compress::Raw::Zlib`. The 36 MB worksheet part goes from 0.174 s to 0.067 s and the 8 MB shared-string part from 0.052 s to 0.023 s. Reading the whole workbook: 0.481 s to 0.323 s, with peak RSS unchanged at 179 MB. - Most of that gap is one thing, and it is not the inflate. `Unzip.pm`'s `ckParams()` sets `crc32 => 1` unconditionally ("unzip always needs crc32"), so every byte goes through `Compress::Raw::Zlib::crc32()` -- 0.076 s on that part, more than the 0.052 s the inflate itself costs -- and the comparison against the stored CRC happens only under `Strict`, which defaults to 0. The check is paid for and never made. This path does not compute it either, which is the same answer for the same money. Asking `Inflate` for `-CRC32 => 1` costs the identical 0.076 s, being the same per-chunk call, so turning it on here would only be worth anything with a croak on mismatch behind it -- and that is a decision about what `read_table` should do with a damaged workbook, not a speed one. - `-Bufsize => 1<<16`: 0.052 s against 0.063 s at the 4 KB default on that part, and flat from there (1<<18, 1<<20 and 1<<22 all measure 0.052 s). - What it declines, handing the archive back to `IO::Uncompress::Unzip` unchanged: zip64 (whether by the end-of-central-directory sentinels or a member's own), split archives, encrypted members, a compression method that is neither stored nor deflate, a directory that does not check out against its own signatures and lengths, a local header that is not one, and an inflate that does not end where the directory says it should. Declining is not an error -- it reads the central directory and one deflate stream and leaves everything past that to the module that has been ported to all of it. - Two things worth knowing about the shape of an archive. `IO::Compress::Zip` with `Append => 1` does not extend one: it writes a second complete archive after the first, and the final end-of-central-directory record describes only that second one, with offsets counted from where it begins. It reads as a single archive only to a reader that scans local headers forwards, which is what `IO::Uncompress::Unzip` does -- so `t/read_table.xlsx.t` and `t/read_table.xlsx.parser.t`, whose fixtures are built that way, go on exercising the fallback. Real workbooks do not have this: Excel, LibreOffice, openpyxl and this module's own `write_table` all take the fast path, which `t/unzip_member.t` pins by reading one that `write_table` wrote. Second, `IO::Uncompress` defaults to `Transparent => 1`, so a file that is not an archive at all comes back as its own raw bytes rather than failing. That has always been true of `_unzip_member`; the fast path declines such a file and changes none of it. - `Compress::Raw::Zlib` joins the prerequisites. It is core since 5.8.8 and `IO::Uncompress::Unzip` was already loading it, but it is dual-life and it is now named directly. - It also reads members the fallback cannot. `IO::Uncompress::Unzip` 2.020 (perl 5.10.1) and 2.024 (5.12.5) refuse a stored member written in streaming mode -- "Header Error: Streamed Stored content not supported" -- and fail to find any member past the first in a streamed multi-member archive at all, both of which this path reads on every perl in the matrix. That is why the new tests assert an accepted member against the bytes that went into the fixture rather than against what the fallback makes of it, and cross-check the fallback only where it can answer. - `t/unzip_member.t` is new, 82 tests: deflated, stored, empty and streamed members, and one whose sizes are in its local header rather than a trailing descriptor; a name another name begins with and one another ends with; an archive comment, and a comment holding the end-of-central-directory signature as a decoy for the backwards scan; an absent member (answered "handled, no such member", so that the sharedStrings.xml plenty of workbooks lack does not pay for a second scan); zip64, a concatenated archive, a file that is not an archive, one that does not exist, an empty one and a truncated one -- each asserted to be declined and then to come back from the fallback exactly as it did before, undef included; and read_table over two workbooks whose directories are intact, one from `write_table` and one with a shared-string table. [A malformed .xlsx could ask for an unbounded amount of memory] - `xlsx_ref_col()` refuses a column reference past XFD, the last of the 16,384 columns ECMA-376 allows, and its comment says why: every row in a sheet is padded to the widest column the sheet mentions, so "a bad one should cost the 16,384 the format allows rather than the twelve million `ZZZZZ` would ask for". It did not. A reference it refuses falls back to "the next column", and that counter had no ceiling at all, so `ZZZZZ1` repeated 20,000 times in one row got 20,000 columns -- the exact outcome the cap was written to prevent, reached through the cap's own fallback. - What it cost: a 54 KB workbook of 20,000 such cells, plus 200 ordinary one-cell rows, came back as 264 MB of empty strings in 0.41 s; 60,000 cells made it 804 MB and 1.35 s, and nothing bounded it but the size of the input. The amplification is about 5,000x, so a 16 MB file of that shape would ask for tens of gigabytes. A cell with no `r=` at all reaches the same counter, so it did not even need a malformed reference. - `xlsx_ws_scan()` now clamps the column index to XLSX_MAX_COL on that path too. Past the ceiling the cells pile up in the last column, last one winning, which is what a repeated `r=` in one row already did. The same line runs on both passes, so they go on agreeing about the width. The two cases above now stop at 16,384 columns and 234 MB, and stop growing: 60,000 cells cost what 20,000 do. - `t/read_table.xlsx.parser.t` covers both routes to the counter -- 20,000 unreadable references and 20,000 cells with no `r=` -- asserted on `_parse_xlsx_sheet_xs` directly, because `read_table` would fold the unnamed columns into one key and hide the width. Both fail without the clamp. [read_table on .xlsx: 42 MB less peak RSS] - `_unzip_member()` returned the decompressed part as a string, which costs a full copy of it: perl cannot hand back a lexical's pad slot, so `return $content` copies. On the 36 MB worksheet part of the 21,845 x 50 workbook that was 34 MB of peak RSS for nothing -- 82.4 MB against 48.5 MB for a scalar reference, and 81 MB still resident after the string went out of scope against 13 MB, the rest being heap the allocator never gave back. - It returns a reference now and its four callers dereference. `$$ws` on an argument list pushes the SV itself, so the part still reaches `_parse_xlsx_sheet_xs` and `_xlsx_sst_xs` without a copy. Reading the whole workbook went from 220.2 MB and 0.487 s to 177.9 MB and 0.475 s. [The `autodie` dependency is gone] - `autodie` was a prerequisite for five calls in one file: three `close`s and two `open`s in `lib/Stats/LikeR.pm`, all of them reading this module's own POD or peeking at the first line of a CSV. Nothing in `LikeR.xs` used it -- the XS side opens through `PerlIO_open` and croaks for itself -- and no test loaded it. `open`/`close` now check for themselves through `_open_read()` and `_close()`, and the pragma and the prereq are both dropped. - Those two helpers raise the failure at the same points, with the message `autodie::exception` 2.37 would have built: `_format_open` (through `_FORMAT_OPEN` and `_format_open_with_mode`) for the open, `_format_close` for the close, each followed by `add_file_and_line`, which is why `_io_die()` takes the file and line one frame up rather than its own. The text is byte-identical, checked against a script running under `autodie`. The one visible difference is that `$@` now holds a plain string instead of an `autodie::exception` object, which nothing here ever inspected. - Two error paths the pragma had been masking are now visible, and both were dead code: `_pod_open`'s `open ... or return undef` and `read_table`'s `open ... or die "read_table: can't open $file: $!"` could never run, because `autodie`'s `open` died first. They are removed rather than revived, so a failed open stays fatal and keeps the wording it had. [Prerequisites: one added, three moved to the test phase] - `IO::Uncompress::Unzip` is now declared. `_unzip_member` has required it for every `.xlsx` read since 0.24 and it was never in the list; it is core since 5.9.4, so the 5.010 floor already guaranteed it, but it is dual-life and a used module belongs in the metadata. `Cwd` stays for the same reason -- `provenance_path()` in `LikeR.xs` reaches it through `load_module()`, which croaks rather than degrades if the require fails, so the tex and xlsx writers genuinely need it. - `Test::Exception`, `Test::LeakTrace` and `Test::More` move from `[Prereqs]` to `[Prereqs / TestRequires]`. They had been runtime requires since 0.18, which made every user of `mean()` install two modules that are not core and that nothing under `lib/` or `LikeR.xs` loads. - They were flattened into the runtime phase by 0.18's "fix to dist.ini for dependencies", but the smoker failure that fix was chasing was not theirs. 0.17 had put `Devel::Confess` in `[Prereqs / DevelopRequires]`, a phase no CPAN client installs, while `t/col2col.t` and `t/transpose.t` still said `use Devel::Confess` -- so the test files would not compile on a smoker. 0.18 cured it by moving every phase back into runtime, which swept the three `Test::` modules along; 0.22 then re-declared `Devel::Confess` as a runtime dependency for the same reason, and 0.24 fixed it properly by dropping the module from the tests. Nothing under `t/` has loaded it since -- both files mention it only in comments -- and every other module the tests use is either core at 5.010 or declared. - Checked, not assumed. `dzil build` puts the three under `test.requires` in `META.json`; `ExtUtils::MakeMaker` carries them as `TEST_REQUIRES`; the generated `Makefile.PL`'s `%FallbackPrereqs` block folds them back into `PREREQ_PM` on any `ExtUtils::MakeMaker` older than 6.63_03, which was confirmed by running the configure step with `$ExtUtils::MakeMaker::VERSION` forced to 6.48 -- an old toolchain sees exactly what 0.316 gave it. `cpanm --showdeps` on the built tarball lists all three. The built dist configures, compiles and passes its 154 files clean. - The repo's own `Makefile.PL` is excluded from the dist and so is never regenerated by dzil; its prerequisite list still named `Devel::Confess` and `Digest::SHA`, neither of which has been used for many releases. It now mirrors `dist.ini`. [Tests] - `t/read_table.xlsx.parser.t` is new: the tokenizer's own cases, which a well-formed workbook never reaches. Cells with no `r=` and with `r=` not first, self-closing `` and ``, blank rows, mid-row gaps, every entity form and the ones that must not be decoded, a shared-string index that is out of range or not a number, `t="str"`/`t="b"`/`t="e"`, a `>` inside an attribute value, `XFD` and the references past it, a repeated reference in one row, a cell with no ``, an empty sheet, a header with no data rows, and the fast path and the callback path agreeing cell for cell. Leak checks on every path through the row emitter, with no `qr//` inside a measured block. - Those leak checks measure the parser with the worksheet part already decompressed. `read_table` itself is measured too, but only where decompressing is clean: `perl-5.10.1`'s bundled `Compress::Raw::Zlib` leaks 18 SVs per read in `crc32(undef)` (`IO/Uncompress/Unzip.pm` line 608), and a check around `read_table` counts those as this module's. The test asks this perl whether it leaks rather than naming versions, and skips only the four checks that need a file. - The whole suite passes on all seven perls in the local matrix -- 5.10.1, 5.12.5 (`long double`), 5.42.3-thr, 5.44.0, 5.44.0+x87, 5.44.0-quadmath (`__float128`) and 5.44.0-i686 (`ivsize=4`) -- and the generated `.c` compiles clean under strict `-std=c99`. Valgrind reports no errors over the truncation and mutation corpora. [aov() could abort the interpreter, or answer differently on every run] - `aov` built the two halves of an `a:b` interaction in two 256-byte stack arrays, and filled the left one with `strncpy(left, term, colon - term)`. `strncpy` writes exactly the count it is given and knows nothing about the destination, so an interaction whose left component ran past 255 characters wrote off the end of the frame, and the `left[colon - term] = '\0'` after it stored past the end as well. glibc caught it as `*** buffer overflow detected ***` and aborted the interpreter, which no `eval` can catch. Both halves are now copied to the heap at the length they actually have. (`right` had been moved off `strcpy` already, but only as far as a truncating `snprintf`.) - Only the FIRST value of a hash-of-hashes was checked for being a reference. `SvRV()` on a later plain scalar reads a pointer out of a field that does not hold one, so `aov({r1 => {...}, bad => 42}, 'y ~ g')` segfaulted -- or did not, depending on hash order, since a non-reference that came first was rejected by the shape test. `lm` and `glm` have always checked this in `lm_read_rows`; `aov` now makes the same check with the same message. - The row count for a hash-of-arrays came from whichever column `hv_iternext()` returned first, so on a ragged frame which observations were fitted moved with perl's hash order from one run to the next. Refused now, naming the column, exactly as `lm_read_rows` refuses it. - `group.stats` had the same defect on the reporting side, and reached it by the documented no-formula form -- R's `stack()` -- whose columns are unequal by construction. `aov({short => [1..3], long => [1..20]})` reported `long`'s mean as 10.5 and its n as 20 on some runs and as 2 and 3 on others, on the same data in the same process image. Each column is now summarised over its own length, which is also what stops the evaluation running past the end of the short ones. - The groups of that no-formula form are stacked in sorted key order rather than hash order. The F statistic survived the shuffling, being invariant to the order of the rows, but only to within rounding -- `Pr(>F)` came back 0.0363396692989842 on one run and ...43 on the next -- and `fitted.values` named entirely different rows each time. [The `.` in a formula expanded in hash order] - `get_all_columns()` returned the column names in hash-iteration order, which perl randomises per process and per hash. `.` is what fixes the term list, and a sequential (Type I) sum of squares is attributed in term order, so `y ~ .` produced a different ANOVA table on every run of the same script over the same data: over four runs of `aov({y, g, x, z}, 'y ~ .')` the `z` row came back with a sum of squares of 15 once and of 1.9e-30 another time, because `z` had been entered before or after the term it is orthogonal to. `lm` and `glm` take their `.` expansion through the same function, so their coefficient order moved the same way. - R has a column order to expand in and a Perl hash has none, so the names are sorted. That is the only order available that is the same twice. [aov()'s formula was truncated, and read its intercept markers with strstr()] - The formula was copied into a 512-byte stack array and stopped at 511 characters, so a longer one was silently truncated and a different model was fitted than the one asked for. `.` expanded into a 2048-byte array and simply DROPPED every column that no longer fit, so `y ~ .` on a wide frame quietly fitted a smaller model. - `-1`, `+0`, `+1` and a leading `1+` were removed with `strstr()` over the whole right-hand side, so the `-1` inside `I(x-1)` was eaten and `y ~ I(x)` -- a different model -- was fitted and reported under that name. - All of it now goes through `lm_formula_split()`, which is what `lm` and `glm` already parse with: it steps over `I(...)`, grows with the formula, and its buffer is on the save stack so a croak anywhere below releases it. The `.` expansion grows through `lm_append()`. [interpolate()'s splines solved a banded system densely] - `ip_build_cubic()` and `ip_build_quad()` assembled an n x n dense matrix and ran a dense Gaussian elimination over it, where n is the number of numeric anchors in the column. Neither system is dense: the not-a-knot cubic spline is tridiagonal apart from its two boundary rows, and the degree-2 B-spline collocation matrix has three non-zero basis functions per row. Both are `kl = ku = 2`. - `ip_solve_band()` solves them in band storage, which is 7 NVs a row. A column of 12,800 anchors went from 0.835 s and 574 MB to 0.002 s and 13.6 MB; one of 400,000 -- which wanted 1.28 TB and could only ever fail -- now takes 0.071 s and 115 MB. `ip_eval_quad()` sums only the basis functions whose support reaches the point, rather than all n of them, so filling g gaps is O(g) and not O(n*g). - The band width is asserted rather than assumed: the degree-2 basis is a partition of unity, so every row of the collocation matrix sums to 1, and a window that had missed a non-zero basis function could not. Checked against SciPy's `CubicSpline(bc_type='not-a-knot')` and `make_interp_spline(k=2)` at 200, 2,000 and 20,000 anchors; worst relative disagreement 2.3e-16. [rbinom() cost O(size) per variate] - `generate_binomial()` was the textbook Bernoulli loop: `size` draws from `Drand01()` per variate, counting successes. Exact, and unusable at any interesting size -- `rbinom(n => 10, size => 1e9)` asks it for ten billion uniforms, and `rbinom(n => 1e4, size => 1e5)` for a billion. 2000 variates at `size => 1e7` took 74 s. - Replaced by BTPE (Kachitvichyanukul and Schmeiser 1988, CACM 31, 216-222), transcribed from R 4.6.1 `src/nmath/rbinom.c`: the inverse-CDF walk below `n*p = 30` and the triangle / parallelogram / exponential-tail rejection above it. Neither draws a number of uniforms that grows with `size`. The same 2000 variates now take 0.0001 s, at any size. - Two departures from upstream. R caches the setup in file-static globals and its own comment there reads "FIXME: These should become THREAD_specific globals"; every variate of one `rbinom` call shares its n and p, so the setup is computed once per call into a caller-owned struct instead -- the same saving, no statics. And `unif_rand()` is `Drand01()`, which is what every other draw in this file uses and what makes `srand($seed)` govern the result. - THE SEEDED STREAM MOVED. BTPE consumes a different number of uniforms per variate than the Bernoulli loop did, so a given `srand` produces different numbers from those 0.315 gave. The distribution is unchanged and a run is still reproducible; only a script that hardcoded the values one seed used to give will see new ones. [ks_test(exact => 1) had no cost cap on the one-sample branch] - `K2x()` builds an m x m matrix, `m = 2*floor(n*D) + 1`, and raises it to the n-th power: O(m^3 log n) time and O(m^2) memory, with nothing but the statistic bounding m. The default route cannot reach a large m, taking the exact branch only below n = 100, but `exact => 1` could. A sample of 800 whose D is 1 -- any badly-fitting reference distribution -- took 52 s and 69 MB, 3200 would have taken most of an hour, and past n ~ 23,000 the cell count overflowed the `int` it was computed in and went to `calloc()` wrapped. The two-sample branch has had `KS_EXACT_MAX_PRODUCT` for this all along. - m is now capped at `KS_EXACT_MAX_M` (500, about half a second and two megabytes) with the same warning and the same fall back to the asymptotic p-value the two-sample branch gives, and `K2x()`/`m_power()` compute in `size_t`. [dnorm() was computed at a double's width on whatever perl it ran on] - `c_dnorm()` decides where the density has underflowed from the exponent range of the floating-point type, and asked `` about a *double* -- `DBL_MAX`, `DBL_MIN_EXP`, `DBL_MANT_DIG` -- whatever perl's NV was. On a long-double or `__float128` build that cut the tail off at |x| ~ 38.57, where a double's subnormals run out, and returned a flat 0 beyond it: `dnorm(-100)` is 1.4e-2174, four thousand orders of magnitude inside a quadmath NV's range, and came back 0. `NV_MAX` / `NV_MIN_EXP` / `NV_MANT_DIG` now, which are the same constants on a double build. - `NV_MIN_EXP` needed a fallback: perl.h has defined it since 5.22, and both `perl-5.10.1` and `perl-5.12.5` are in the matrix. `ppport.h` does not backport a macro that is not an API function, so `LikeR.xs` derives it the way perl.h does -- from the same `` constants and the same `USE_QUADMATH` / `USE_LONG_DOUBLE` tests the `nv_*` libm layer already switches on -- under an `#ifndef` a perl that has it never reaches. `NV_MANT_DIG` is guarded beside it so that a build with one and not the other cannot fail pointing at the wrong line. Caught by `./test.all.perls.pl`, which is what it is for: both builds failed at `make`, and nothing on the default perl would ever have shown it. [Convergence thresholds written against a double] - The continued fraction for the incomplete beta, and the series and continued fraction for the incomplete gamma, stopped as soon as a term fell below a bare 1e-15 or 3e-15 -- so `pt`, `pf`, `pchisq`, `qchisq`, `qf` and everything built on them returned about sixteen digits on a perl carrying nineteen or thirty-four. `LIKER_EPS_SCALE` carries each of them to the build's own width; it is `NV_EPSILON / DBL_EPSILON`, which is exactly 1.0 when NV is a double, so every value produced on a double build is unchanged to the last bit. Thresholds that are part of an algorithm's definition -- `FT_TOL`, which is R's `uniroot()` default, and MASS's `double.eps^0.25` in `nb_theta_ml()` -- are numbers from the reference implementation and stay put; each says so where it is defined. - `FT_EPS` looks like one of those and is not: it is the other half of R's `uniroot()` stopping rule, `2*FT_EPS*|b| + FT_TOL/2`, and it is also the lower endpoint of the bracket `fisher_test`'s confidence interval is inverted over -- so it sets the largest odds ratio that interval can name, which R reports as `1/DBL_EPSILON`. Scaling it to the build's own epsilon moved the conditional odds ratio for SciPy's gh-3014 table 1.2e-9 off R on a `__float128` build and turned that table's upper limit from 4503599627370496 into 5.2e+33. It stays `DBL_EPSILON` on every build, and now says why. - `_qgamma()` in `LikeR.pm` inverted `1 - _igamc($shape, $x)`, which is the cancellation `igam()` was added to avoid: below a lower tail of about `NV_EPSILON` that difference can only be a multiple of `NV_EPSILON`, and below about 1e-16 it is exactly 0, so a bisection against it has nothing to bisect on. It now bisects against the new private `_pgamma_lower`, which is `igam()` itself. `age_standardize()`'s Fay-Feuer interval asks for the `alpha/2` quantile, so the lower limit is what this reaches: at `conf.level => 1 - 2e-12` it moved from 1.1e-6 off R to 3.3e-7 off. It cannot be pushed much further from the Perl side -- `1 - 2e-17` is already 1 in a double -- which is why `t/age_standardize.t` checks the primitive itself against R's `pgamma` as well as the interval. [The rank test in aov()'s QR was not scale invariant] - `apply_householder_aov()` declared a column aliased on the absolute `max_val < 1e-10`, which is not a statement about collinearity but about units. A design whose columns are all smaller than 1e-10 -- a predictor in metres that wanted micrometres, a rate per person-year -- had every column declared aliased at step 0, and `aov` reported zero degrees of freedom and a zero sum of squares for every term on perfectly well-conditioned data: `aov` on `x` gave R's F of 1.9927680012954 and on `x * 1e-12` gave NaN. The test is now relative to each column's own scale, taken before the reduction starts, which is what `sweep_matrix_ols()` has always done for `lm`. [strtok() in the formula parsers] - `lm_formula_terms()` and `aov` split the right-hand side with `strtok()`, whose position lives in a libc static. On a `-Dusethreads` perl every thread is its own interpreter inside one process and shares that static with all the others, so two threads fitting a model at the same moment could each be handed the other's term list. Replaced by `lm_tok()`, which takes the cursor from the caller and has no state of its own. It is the only routine with hidden state this file used. [Five symbols the shared object should not have exported] - `approx_pnorm()`, `igamc()`, `get_p_value()` and the `cs_uninit_catcher` XSUB were compiled with external linkage, so the `.so` exported them alongside `boot_Stats__LikeR()`. `igamc` in particular is a name a numerical library might well define too, and the dynamic loader resolves the first definition it sees -- an interposed one would silently replace every chi-square tail this module computes. All four are static now. A fifth, `compare_doubles()`, had had no caller since the `qsort()` comparators were replaced by `LIKER_DEFINE_SORT()`, and is gone. - The generated `.c` is also clean under `-Wsign-compare`, which is not in `Makefile.PL`'s flags but is the check the mixed-sign comparisons this file's type rules can introduce would show up in. Seventeen of them were left; all were benign, and all are gone. [Smaller: memory and time that was being spent on nothing] - `evaluate_term()` `savepv()`d the term string -- a malloc, a strcpy and a free -- for every cell of every design matrix `lm`, `glm`, `aov` and `anova` build, in order that two branches it was not going to take could write NULs into it. A bare column name, which is what nearly every cell asks for, now allocates nothing. - `melt()` built the whole long frame as one throwaway record per output row, with a nested arrayref of the id values, before materialising it: a melt of R rows over V value columns held R*V of them alive at once beside the result they were about to become. It emits into the requested shape as the loops go. - `table_one()` built its per-group row lists with one `grep` over the whole frame per group -- O(groups x rows), the same shape as the O(levels x groups x rows) counting beside it that 0.315 replaced with a single pass. One bucketing pass now. - `csort()`'s AoH -> HoA materialisation walked the sorted rows once per output column, fetching and type-checking each row `nk` times for the n*nk cells it produces. Row-outer now, one fetch per row, columns allocated at their final length and filled through `AvARRAY` as `filter()` and `mg_column()` do. - `colnames()`, `_present_keys()` and `_rename_inplace()` flattened a copy of every row reference in the frame (`my @rows = @$df`) to read it once; `assign()`'s HoA row view is filled with one hash slice rather than a keyed store per column. [Tests] - `t/aov.regressions.t` pins every one of the `aov` items above: the long interaction component and the long formula (the answer must not depend on how long a column's name is), the malformed HoH, the ragged HoA, the determinism of `group.stats` and of `.`, the scale invariance of the rank test, and `I(x-1)`. Against the code these entries replace it fails in thirteen places, and the `ks_test`, `rbinom` and `interpolate` files below hang or exhaust memory on it rather than merely failing. - `t/interpolate.spline.banded.t` checks the banded solve against SciPy at 200, 2,000 and 20,000 anchors, with the generator committed beside it as `t/interpolate.spline.banded.py`, and interpolates a column of 200,000 -- which the dense build cannot allocate. - `t/rbinom.dist.t` runs a chi-square goodness-of-fit against the exact binomial CDF over thirteen `(size, prob)` pairs chosen to cross both branches of BTPE and the reflection, plus the moments, the support, the short circuits and reproducibility under `srand`. It deliberately pins no individual variate: pinning one would pin the algorithm rather than the distribution. - `t/ks_test.exact.guard.t` pins the one- and two-sample exact p-values against R and asserts that a forced exact run past the cap warns and falls back; `t/value_counts.utf8.t` checks `value_counts` against what a Perl hash makes of the same list, in both directions -- "\x{e9}" and "\xe9" are one value, "\x{263A}" and its three bytes are two; `t/dnorm.nv_width.t` pins `dnorm` against R and asserts the identity `dnorm(x) == exp(-x^2/2) / sqrt(2*pi)` at whatever width the perl running it carries -- which is one assertion that covers every build, since both sides underflow together on a double and neither does on a wider NV. Its R table stops at |x| = 29 on purpose: `perl-5.10.1` reads `2.1200065515246056e-298` as `1.999999999999999e-298` and anything below ~1e-308 as 0, so a table of R's far-tail values would be testing perl's own `atof`. - `t/age_standardize.t` gains the gamma quantile at a small tail probability, pinned against R's `qgamma` at four confidence levels, and checks `_pgamma_lower` directly against R's `pgamma` -- `conf.level` cannot reach far enough into the tail on its own to separate it from `1 - _igamc`, because `1 - 2e-17` is already 1 in a double. Where that subtraction starts losing the tail is a build property and the test asks the perl for it rather than assuming a double: its first draft pinned a literal 1e-20 and passed everywhere except `__float128`, where 1e-20 is still fourteen orders of magnitude above the point the subtraction fails at. - All of it passes on all seven perls in the local matrix -- 5.10.1, 5.12.5 (`long double`), 5.42.3-thr, 5.44.0, 5.44.0+x87, 5.44.0-quadmath (`__float128`) and 5.44.0-i686 (`ivsize=4`) -- 41,166 tests on the wider NV widths and 41,160 on the rest, the difference being the six `dnorm` assertions that only a build whose exponent range reaches past a double's has anything to check. The generated `.c` compiles clean under strict `-std=c99` and under `-Wall -Wextra -Wsign-compare`, on a quadmath CORE as well as a double one. 0.3151 2026-09-05 CDT [Two leak checks failed on perl 5.10.0, which leaks every `qr//` itself] - A CPAN smoker running perl 5.10.0 (x86_64-linux, `double` NV, 64-bit `IV`) failed 0.315 in `t/cfilter.t` and `t/filter_match.t`, one leaked SV each. `Test::LeakTrace` reported a `PV` holding the string `"Regexp"` -- `CUR = 6`, `REFCNT = 1` -- and attributed it to the line of the `no_leaks_ok` call. Nothing in `LikeR.xs`, and nothing on the Perl side, allocates that SV. - It is perl's own. `pp_qr()` blesses the object it builds into the package named by `reg_qr_package()`, which is a `newSVpvs("Regexp")`; on 5.10.0 nothing ever releases it, so *every* evaluation of a `qr//` leaks one SV. The `SvREFCNT_dec(pkg)` that fixes it is in 5.10.1's `pp_hot.c` (and `reg_qr_package()` is `regcomp.c:5287` there). `perl-5.10.1` is the oldest perl in the local matrix, so no run here could reproduce this, and none of the newer perls the smokers use can either. - Those two blocks were the only two leak checks in the whole suite that evaluated a `qr//` inside the measured block, which is why exactly two subtests failed and why the count was exactly one SV apiece. Both now compile the pattern into a lexical ahead of the block and pass that in. `cfilter`'s regex selector and `col()->match` each take a precompiled `qr//` as it comes -- `col()->match` compiles a pattern with `qr//` only when handed a string -- so the code under measurement is unchanged, and the checks still cover the same paths. - Nothing in the module changed: 0.315 does not leak on perl 5.10.0, and `t/cfilter.t` and `t/filter_match.t` are the only files that differ. 0.315 2026-09-04 CDT - Defect fixes, plus a set of speed and memory changes. No numeric answer moves anywhere in the release except `cov`'s on an incomplete pair, which disagreed with R and is called out below; every reorganised routine was checked against the code it replaced and returns the same bits. - The defects came from two sweeps, which is why there are two groups of segmentation faults below. - The first sweep was fuzzing rather than reading: every exported function called with a reference where it expected a number, with a code ref or a scalar ref buried inside an otherwise ordinary frame, and with randomly assembled argument lists, all run against an AddressSanitizer/UBSan build. The whole existing suite is clean under those sanitizers and always was; what the fuzzers reached is the argument space the suite does not cover, because it is the space no correct program visits. A user's typo visits it. - The second was the argument-validation surface again, under valgrind with `--errors-for-leak-kinds=definite` and with RSS watched over tens of thousands of failing calls. Fuzzing finds the call that dies; a leak on a croak path does not die, so nothing but the allocator notices it. Six functions were leaking, and seven could be made to crash or hang. [Six segmentation faults and two runaways] - `av_fetch()` returns NULL for a hole in a sparse array -- what `delete $a[2]` leaves behind, and what `$a[0] = 1; $a[3] = 4` makes of indices 1 and 2 -- and four readers dereferenced it without looking. None of it could be caught: `eval` does not see a signal. - | call | before | now | |---|---|---| | `epi_2x2([10, 20, , 40])` | `SIGSEGV` | croaks `cell at index 2 is undef` | | `epi_2x2([[1, 2], [, 4]])` | `SIGSEGV` | croaks `row 1 cell at index 0 is undef` | | `cmh_test([[10, 20, , 40]])` | `SIGSEGV` | the same reader, same message | | `binom_test([, 5])` | `SIGSEGV` | croaks `successes is undef` | | `fisher_test([[1, 2], [3, ]])` | `SIGSEGV` | croaks `array cell is undef` | | `coxph(\@t, \@s, \@x)`, hole in any | `SIGSEGV` | croaks, naming the vector and index | - `survfit` and `logrank_test` had the same fault with a different ending. `srv_read()` did test the pointer, but mapped the hole to `NaN`, and a `NaN` survival time does not fail the `t < 0` check that follows -- every comparison against one is false -- so it reached the risk-set construction and the fit ran away instead of returning. A time that is not a number is refused outright now. - All of it goes through one reader, `av_num_at()`, which names the offending index in the message: a hole is invisible at the call site, since nothing about `[10, 20, 40]` printed from a sparse array says which slot is missing. - That reader also rejects a cell that is not a number, where the old code read it as zero. `epi_2x2([1, 2, 'abc', 4])` used to return an odds ratio computed from a silent 0 and now croaks. This is what R does -- `chisq.test`, `fisher.test` and `binom.test` all refuse to coerce -- and what this file's own `ct_cell()`, `ft_cell()` and `bt_check_count()` already did. [Six leaks on croak paths] - `croak()` longjmps past any `Safefree()` written after it, so a validation failure that fired once the working set was allocated dropped all of it. Found by running the suite under valgrind and by watching RSS over tens of thousands of failing calls; per failing call, before: - | call | leaked | |---|---| | `dunn_test(..., method => 'nope')` | 113 KB at n = 2000, growing with n | | `scale([...])` with a non-numeric element | 8 bytes per element (160 KB at n = 20000) | | `aov($d, 'y ~ a:b')` with no main effects | 7.7 KB | | `col2col($d, 'sum', 'nosuchcolumn')` | 3.7 KB, growing with rows x columns | | `fisher_test` with a non-numeric cell | 3.2 KB, growing with the table | | `glm` with a response outside its family | 2.9 KB | - Each was fixed at the level it belonged to rather than by one blanket rule. `dunn_test` now settles its method before the first allocation, which is where `cov()` and `p_adjust()` already ask the same question -- it used to be `dunn_padjust()` that noticed, from the bottom of its dispatch chain, with the values, the group labels, the ranks and five p-value arrays all outstanding. `glm`'s three response checks moved above the IRLS working set; all three read nothing but the response, which does not change once the design is built, so asking once is exactly what asking per observation per outer pass did. `aov`'s interaction croak got the free list it was missing. `scale`, `fisher_test`, `coxph` and `srv_read` put their buffers on the save stack, as `wilcox_test` already did. - `col2col` needed more than a free list. With `skip.errors` turned off, a caller's block that dies propagates out from the middle of the per-pair loop, past every `Safefree()` in the function, so no hand-written cleanup could cover it; the column tables go on the save stack instead. Separately, none of that function's sixteen croaks released the column-name array, so each also leaked an AV and one SV per column -- it is mortal now. [`cov` now agrees with R on an incomplete pair] - This one changes an answer. `cov` compacted to the pairwise-complete observations, so `cov([1,2,undef,4,5], [1,2,3,4,6])` was 4. R's `cov` defaults to `use = "everything"` and propagates, so R 4.6.1 gives `NA` there, and `cor` in this same distribution already gave `NaN`. `cov` now gives `NaN` too. - `cor_test` still drops incomplete cases, because R's `cor.test` runs `complete.cases()`. So all three now match their R counterparts, where before `cov` was the one that did not. There is no `use` argument to ask for the old behaviour with; drop the pairs before the call to get it. - A comment in `LikeR.xs` claimed `cor` dropped pairs the way `cov` did. It never has. Corrected. [`cor` on a matrix is up to 24x faster, and holds less] - `cor($matrix, undef, 'spearman')` ranked both columns of every pair, so a p-column frame paid p(p-1) rankings where p answer the whole question. Each column is ranked once now and the pairs are Pearson on those ranks, which is the same arithmetic: a column's ranks do not depend on its partner, because `cor` does no pairwise deletion. On 500 rows, 0.0051s to 0.0005s at p = 20, and 0.2955s to 0.0125s at p = 160. Pearson and Kendall are unaffected, and the method string is now resolved once instead of at each of the p(p-1)/2 pairs. - `cor($matrix)` built an ncols x ncols scratch matrix of correlations only to copy it cell by cell into the result and free it. Each value is stored at both of its positions directly. Peak memory on a 600-column frame fell from 32.1 MB to 29.3 MB. [`table_one` no longer rescans the frame once per level] - Categorical counts came from a `grep` over a group's rows for each (level, group) pair, run twice over -- once for the table the test statistic sees, again to format the output rows -- plus one scan of every row per level for the Overall column. One pass per group now feeds a lookup. On 20000 rows, 0.021s to 0.010s at 5 levels and 0.124s to 0.010s at 40, and flat in the level count rather than unbounded in it. [Four segmentation faults and a `SIGFPE`] - Each of these ended the interpreter. None of them could be caught: `eval` does not see a signal, so a caller could not even fail gracefully. - | call | before | now | |---|---|---| | `matrix([1..6], 0)` | `SIGFPE` | croaks `Dimensions must be greater than 0` | | `matrix([1..6], 'greater')` | `SIGFPE` | croaks `matrix: nrow must be a number` | | `prcomp({ c => {d=>1}, e => undef })` | `SIGSEGV` | croaks `prcomp: HoH value for row 'e' is not a hash-ref` | | `scale([[1,2], sub {1}])` | `SIGSEGV` | the row reads as absent, as a string or `undef` row always did | | `merge(\1, [{a=>1}])` | `SIGSEGV` | croaks `merge: left frame must be AoH/HoA/HoH` | - `matrix`. The guard against a zero dimension sat one branch *after* the inference that divides by it, so `ncol = (data_len + nrow - 1) / nrow` ran first with `nrow == 0`. A non-numeric string reached the same division because `SvUV('greater')` is `0`. The guard now runs before the inference, and both dimensions go through the validating reader below. - `prcomp`. A HoH frame's shape is decided from whichever row `hv_iternext` reaches first, and nothing checked the rest — both the column-name pass and the extraction pass called `SvRV` on every value and dereferenced the result as an `HV`. HoA and AoA already had a rectangularity pre-pass; HoH now has the same one. Because the deciding row moves with hash order, the crash came and went between runs on identical input, which is the worst way for a bug to present. The HoA branch reached the same null pointer by a second route: column names are copied with `savepv` and looked up again with `strlen`, so a name holding a NUL byte truncates, `hv_fetch` misses, and the `NULL` was dereferenced. That is now a croak naming the column. - `scale`. The per-row test in matrix mode was `SvROK()` alone, without the `SvTYPE(...) == SVt_PVAV` half, so a reference to anything that is not an array was handed to `av_fetch` as an `AV`. Only row 0 is checked when the matrix shape is detected; every other row arrived unvalidated. A bad row now takes the same path a plain string or an `undef` row has always taken — it reads as absent — rather than becoming a new croak, which would have changed two cases that were never broken. - `merge`. `mg_shape` ran on nothing but an `SvROK` test and treated everything that was not an array as a hash, so it reached `hv_iterinit` on a scalar ref *before* `mg_prep` — which is where a frame is really validated — could reject it. It now answers `0` for anything that is neither, and leaves the croak to `mg_prep`, which is where the message the caller should see comes from. [`bw_ucv` and `bw_bcv` could loop forever] - Both build their search interval from `sqrt(var(x))`, which squares the data. On a `double` NV the variance of `c(1, 1e300, -1)` is already `+Inf`, and `dens_brent_fmin` — R's `Brent_fmin`, transcribed — iterates until the bracket closes, which a non-finite bound never does, because every comparison against a NaN is false. `bw_ucv([1, 2, 3, 1e300])` and `density(\@x, bw => 'ucv')` on the same data spun with no way out but a signal. - R does not reach that loop either: `optimize()` rejects the bounds first, with *invalid 'xmin' value*. This is the same check, moved to where the bounds are built. The threshold is a property of the build — about `1e155` for a `double` NV, far higher for long double and `__float128` — so the check tests the value, not the data. - `dens_brent_fmin` also gained an iteration cap as a backstop against a future caller. It is not a working limit: golden-section alone shrinks the bracket by 0.382 a step, so even `tol = NV_EPSILON` on a 113-bit `__float128` needs about 163 steps, and the cap is 1000. [Size arguments are validated in one place] - `matrix`, `hist`, `sample`, `rnorm` and `rbinom` each read a count with a bare `SvUV`/`SvIV`. `SvUV` of a non-numeric string is `0`, of `-1` it is `2**64-1`, and of a reference it is the address — and each of those went straight to a divisor or an allocation: rnorm(-1) # Out of memory in perl:util:safesysmalloc rbinom(n => -1, size => 2, ...) # Out of memory in perl:util:safesysmalloc sample([1,2,3], $some_ref) # Out of memory in perl:util:safesysmalloc hist(\@x, sub { ... }) # Out of memory in perl:util:safesysmalloc - Perl's out-of-memory death is unrecoverable — it is not a croak and `eval` never sees it either. All five now share one reader, `sv_count_arg`, which refuses `undef`, a non-number, a NaN, a negative and anything above `2**48`, and truncates a fraction toward zero the way R's `as.integer` does — so `matrix(1:6, 2.7)` is a two-row matrix, in R and here. `runif` already had a hand-written check for exactly this; its message is unchanged. [`cor_test` reports `NaN` where the correlation is undefined] - A column with no variance has no correlation, and R says so: `cor.test(c(1,1,1,1), c(1,2,3,4))` gives `cor = NA`, `t = NA`, `p-value = NA`, `df = 2`. This returned estimate `0`, statistic `0`, p-value `1` — which reads as a real, well-supported null result that no caller can tell apart from one. Three code paths in this module already disagreed with it: `cor()` croaks (*standard deviation of x is 0*), the shared `pearson_cor` helper returns `NV_NAN`, and `cor_test`'s own Kendall branch already answered `NaN` on its degenerate denominator. Only the Pearson and Spearman branches returned `0`. - All three methods now report `NaN` for estimate, statistic and p-value, and Pearson reports a `NaN` interval and keeps `df = 2`, as R does. Ordinary correlations are bit-identical. [`sample` draws without replacement, and now says so] - Asked for more than the population holds, the two shapes disagreed and neither said anything: sample([1, 2, 3], 10); # [3, 2, 1, undef, undef, undef, undef, undef, undef, undef] sample({a=>1, b=>2}, 5); # two keys, silently - Seven undefs that no caller could distinguish from real data is the worse of the two. Both branches now croak, as R does — *cannot take a sample larger than the population when 'replace = FALSE'*. Draw sequences under a fixed `srand` are unchanged: the shuffle makes exactly the same number of `Drand01()` calls for every call that was previously valid. A first argument that is neither an array nor a hash reference is now a usage croak rather than a silent `undef`. [`hist` follows R's own `breaks` rule, and its extrema are `NV`-wide] - `breaks` went through a bare `SvIV`, so `hist(\@x, -5)` wrapped to a huge `size_t` and silently behaved like `breaks => 1`. The rule is now R's, from `src/library/graphics/R/hist.R`: anything not finite or below 1 is *invalid number of 'breaks'*, and a value above `1e6` warns and clamps, because `pretty()` needs an `n` that fits in an `int`. A fractional value truncates — R's `breaks = 2.7` gives the same four breaks as `breaks = 2`, and so does this. - Separately, `hist` seeded its running extrema from `DBL_MAX` rather than `NV_MAX`. On a long-double or `__float128` build every value above `DBL_MAX` fails `val < min_val`, so the minimum stayed at `1.8e308`. Measured on `5.44.0-quadmath`: - | | before | now | |---|---|---| | `hist([1e400, 1.05e400, 1.1e400])` | breaks `0 5e399 1e400 1.5e400`, counts `0 1 2` | breaks `1e400 1.05e400 1.1e400`, counts `2 1` | - `min()` and `max()` on the same data were always right, which is what made the empty leading bin visible. [Smaller things] - **`write_table` gained `quiet => 1`.** The confirmation line is deliberately unconditional and deliberately coloured, which is right at a terminal and wrong for a script whose stdout is a pipe or a data file — the SGR bytes go out whatever file descriptor 1 is, and the only advice on offer was to capture and strip them. `quiet` silences the line rather than decolouring it, because the coloured form is the contract every format shares. - **`chunk` owns its argument-list message.** `chunk($aref, 2)` fell into a hash assignment and died as perl's *Odd number of elements in hash assignment at LikeR.pm line 1686*, naming neither `chunk` nor the option it wanted. - **`auto.row.names` is documented where it can be found.** The option has existed and been tested since 0.313, but appeared only in the release notes, not in `read_table`'s own options table — so `h('read_table')` did not list it, and the R-written `mtcars.tsv` in this repository looked unreadable. - `aov` allocated its design matrix and the snapshot it keeps for fitted values a row at a time: 2n allocator calls and 2n malloc headers for a matrix `lm` has always taken in one block. They are one block each now. At n = 20000 that is 39,965 fewer allocations, and `aov`'s time relative to `lm` on the same fit went from 1.118x to 1.002x. The block multiply is guarded against wrapping where `size_t` is 32 bits. - `wilcox_test`'s exact confidence interval sorted the pairwise differences and the Walsh averages with `nv_heapsort` where the introsort in the same file is 1.6x to 1.9x faster at those sizes. The exact interval at n = 800 went from 0.1309s to 0.1162s. - `assign(..., map_cell { ... })` on a hash-of-arrays frame assembled the row view before testing whether the cell was `undef`, so a column of `undef` built one view per row and threw every one away unread -- 0.161s of the 0.165s a 20000-row, 64-column pass cost. The test comes first now, and the view is one hash refilled per row rather than a fresh hash each time. That last part is visible to a block that keeps `$_[0]`: it now sees the current row through the kept reference rather than the row it was handed, which is what the AoH and HoH forms have always done, since those hand the block the row hash itself. [Testing] - `t/arg.crash.regressions.t`, 98 assertions. Every crash case runs in a child perl and is judged by how the child exited, because a signal death cannot be caught in the process that provokes it; the `prcomp` cases run 40 fresh interpreters each, since the crash depended on hash order and reproduced within the first two. R provenance for every quoted behaviour (R 4.6.1) is in the file header, and nothing in it needs R, `python3`, NumPy or SciPy at run time. The two `hist` extremum assertions skip on a `double` NV and run for real on quadmath. `t/chunk.t` and `t/write_table.announce.t` gained the cases for the two smaller fixes. - `coxph` had no test at all. `t/coxph.R.t` cross-validates it against R's `survival` 3.8.9 on two corpora: `ovarian`, which that package ships and `?coxph` uses, and a 16-observation set built for the purpose with ties at two event times, because `ovarian` has none and Efron and Breslow are therefore indistinguishable on it. Its values are all small integers, so the tie structure does not move with the NV width. Worst disagreement with `survival` across all 22 quantities and both tie methods is 8.88e-16 on the double builds, 3.64e-16 on long double and `__float128`, and 1.11e-15 on the 32-bit one -- 2 to 5 ulp. `t/coxph.R.R` regenerates the frozen table. - `t/croak.leaks.t` covers every crash and leak above: that each croaks with a message naming what went wrong, that none leaks an SV, and that the ordinary call still returns what it did. - `t/assign.t` gained two checks for the `map_cell` change, one of them a tied sibling column that counts its `FETCH`es, so that a row view built for a cell the block will never see is a failure rather than a slowdown. - The whole suite runs clean under valgrind with `--errors-for-leak-kinds=definite`, which it did not before: six functions leaked. 0.314 2026-09-02 CDT [`read_table` is 2.5× to 3.6× faster] - `read_table` parsed in C and then assembled in perl: `_parse_csv_file` cut each line into fields and handed the row to a closure, once per row, which rebuilt it as a hash one field at a time. On a 300,000 × 5 CSV that closure was 79% of the call — 0.42 s of 0.53 s — against 0.11 s for the parser feeding it. Not the calls themselves: 300,000 calls into an empty sub cost 0.010 s. It was the work inside them, done 1.5 million times. - Once the header is fixed there is nothing in that closure left to decide, so `_parse_csv_file` now takes a plan — which columns, which field each one comes from, and where to put it — and assembles every remaining row itself. The closure still reads the header, because that is where every message `read_table` can produce comes from; it then hands the rest of the file over and is not called again. The field SVs are *moved* into the row hash or the column array rather than copied, and the row buffer is reused, so a cell costs a pointer instead of a `newSVsv()`. - Measured on the `plot.scaling.pl` fixtures at `n = 300,000`, median of 7, pinned to one CPU, with Python and R timed in the same session on the same files: - | | before | after | speedup | Python | R | |---|---|---|---|---|---| | `read_table` (csv, numeric) | 0.499 s | 0.195 s | 2.6× | 0.064 s | 0.603 s | | `read_table` (csv, mixed) | 0.573 s | 0.228 s | 2.5× | 0.288 s | 0.432 s | | `read_table` (tsv, mixed) | 0.572 s | 0.223 s | 2.6× | 0.289 s | 0.401 s | | `read_table` (csv, hoa) | 0.877 s | 0.330 s | 2.7× | 0.061 s | 0.381 s | - At `n = 100,000` the same four are 3.0×, 3.0×, 3.0× and 3.6× faster. - `read_table` now beats R's `read.csv` and `read.delim` on all four panels, having previously lost three of them. Against Python, the mixed panels are the like-for-like pair — both build 300,000 five-key row records, `csv.DictReader` against `output.type => 'aoh'` — and `read_table` went from 2.0× slower to 1.3× faster. The numeric and `hoa` panels are `pandas.read_csv`, which returns four typed NumPy blocks rather than 1.5 million scalars; that gap is the shape of the answer, not the parsing, and no amount of parser work closes it. - Two shapes keep the old path, because both need perl on every row: `output.type => 'hoh'`, which names each row from one of its own columns, and any read given a `filter`. `.xlsx` has a parser of its own and is untouched. Nothing any of them returns changes. - What did get faster on that path is `hoa`, which used to build a row hash it did not need and then read every key back out of it, and which looked each column array up in the result hash — autovivifying it — once per column per row. The column arrays are made once when the header is finalized now, and pushed to directly. [`t/read_table.fast_path.t`, 44 tests pinning the two paths to each other] - Fifteen fixtures — duplicate column names, empty cells, `na.strings`, quoted fields with embedded commas and newlines, header-only, one data row, `auto.row.names`, commented-out headers, TSV, CRLF, single column — each read twice in both `aoh` and `hoa`, once down each path, and required to come back `is_deeply` identical. A no-op `filter` is what forces the closure, and it cannot change the answer: a filter key of 0 is handed the whole row, and `read_table` writes a mutated `$_` back only for keys above 0. The alignment error is checked to be worded identically by both paths, data row number included. `t/parse_read.t` gained 11 more tests, for the plan's own argument validation — each of those croaks guards a C array that would otherwise be indexed with a number that came from perl. [`seq` rewritten against R's `seq.default`] - `seq` was a `(to - from)/by` loop with one fuzz factor and two guards. R's `seq` is not that. `base::seq.default` is a chain of eight special cases — two different fuzz factors, three separate routes to a single value, a rewrite for endpoints whose difference overflows, and a clamp on the last element — and every one of them decides either how many values come back or where the sequence stops. All of them are transcribed now, from R 4.6.1 `src/library/base/R/seq.R` and `seq_colon()` in `src/main/seq.c`. - It matters which R function is being copied, because `seq.default` and the `seq.int()` primitive do not agree. `seq.int` tolerates a slightly negative `(to - from)/by`, allows `100 * INT_MAX` values, and short-circuits `by = ±1` to `from:to`; `seq.default` rejects any negative count, caps at `INT_MAX`, and reaches `from:to` only when `by` is genuinely absent. This follows `seq.default`, because that is what R's `seq()` dispatches to. - Six calls were wrong, three of them silently: - | call | before | now, and in R | |---|---|---| | `seq(0, 1e30, 1)` | `()` | croaks `'by' argument is much too small` | | `seq(NaN, 5)` | `panic: stack_grow() negative count` | croaks `'from' must be a finite number` | | `seq(0, 1, 1e-11)` | `Out of memory during stack extend` | croaks `'by' argument is much too small` | | `seq(-1e308, 1e308, 1e307)` | `()` | 21 values, `-1e308` to `1e308` | | `seq(5, 1)` | croaks `wrong sign in 'by' argument` | `5 4 3 2 1` | | `seq(1e15, 1e15 + 20, 2)` | 11 values | the single value `1e15` | - The empty lists and the panic were one bug: the element count was computed as an NV and cast to `size_t`, and neither `1e30` nor `NaN` nor `Inf` fits in one, so the conversion was undefined. On x86-64 `1e30` became `0` and the call returned nothing at all; `NaN` became `-2**63`, which then reached `EXTEND()`. A count that *did* fit but could not be allocated — `(1 - 0)/1e-11` is 100,000,000,001 values, 800 GB of stack — got as far as trying. R rejects all three the same way, and now so does this: `'by' argument is much too small` above `INT_MAX` values, `'from'`/`'to' must be a finite number` for a non-finite endpoint. - The other three were missing features rather than broken arithmetic: - **An absent `by` is `from:to`, in either direction.** `seq(5, 1)` is `5 4 3 2 1`, the same as R; only an explicit `by` whose sign disagrees with the direction of travel is an error. `from:to` also carries R's *looser* fuzz, so `seq(1, 4.9999999)` is five values where `seq(1, 4.9999999, 1)` is four — R does the same, and for the same reason: `1:4.9999999` adds `1 + FLT_EPSILON` before truncating and the `by` branch adds `1e-10`. - **The last element is pinned to `to`** when the fuzz carried it past, which R has done since 2.9.0. `seq(0, 1, 0.00025 + 5e-16)` used to overshoot 1 — the exact case R's `tests/reg-tests-1b.R` asserts against, under the heading "(Deliberate) overshot in `seq(from, to, by)` because of fuzz". - **Endpoints too close to tell apart collapse to `from`.** When `abs(to - from) / max(abs(to), abs(from))` is below `100 * DBL_EPSILON` the count `(to - from)/by` is noise, and R returns `from` alone. This is why `seq(1e15, 1e15 + 20, 2)` is one value: at that magnitude a `double` cannot distinguish the endpoints well enough for a step of 2 to mean anything. Widen the gap and the sequence returns — `seq(1e15, 1e15 + 200, 2)` is 101 values. - That last threshold is deliberately *not* scaled to the build's `NV_EPSILON`, which is this file's rule for anything epsilon-shaped. R's `100 * .Machine$double.eps` is a third fuzz factor belonging to `seq`'s definition, like the `1e-10` and the `FLT_EPSILON`, rather than a tolerance on this build's arithmetic. Scaling it was tried and reverted: with `NV_EPSILON` the call above returns eleven values on a long-double or `__float128` perl and one on a `double` perl, which is a worse answer than R's on every build. [`seq`: integer sequences come back as integers, and are up to 5× cheaper] - A sequence whose every value is an exact integer no larger than `2**53` is now returned as perl integers (IVs) rather than floats — which is also what R returns for the same call, an integer vector. The numbers are identical either way; this is a choice of SV body, not of arithmetic. But it is much the cheaper body, because stringifying an IV never reaches `Gconvert()`. - Measured on perl-5.44.0 at `n = 1e6`, per element: - | | before | after | perl's own `1 .. n` | |---|---|---|---| | `my @a = seq(1, n)` | 16.3 ns | 12.7 ns | 12.3 ns | | `join ',', seq(1, n)` | 409 ns | 81 ns | 81 ns | - The second row is the one that shows up in real code: anything that prints, joins, writes or hash-keys the result was paying five times over. `seq` is now level with the range operator on both. A fractional step stays floating point, as it must, and is unchanged at 13.4 ns per element — that path was already at parity. - Void and scalar context no longer build the values the caller cannot reach. `seq(1, 1e7);` as a statement of its own took 190 ms and left 384 MB of resident memory behind for the life of the process — the grown argument stack, the grown mortal stack, and the SV arenas the ten million heads came out of. It now takes 5 µs and no memory. In scalar context only the last value is built, which is the one the caller's assignment reads off the top of the stack: the same answer perl gives for any list-returning sub, and the same answer `seq` gave before. - One consequence is cosmetic: a large integral value prints in full rather than in exponent form, so `seq(1e15, 1e15 + 200, 2)` starts `1000000000000000` where it used to start `1e+15`. [`t/seq.R.t`, 378 tests taken from R's own suite] - Every case is R's, cited individually in the file's header: the `## seq` and `## Round` blocks and the Don MacQueen length case from `reg-tests-1a.R`, the deliberate-overshoot assertion from `reg-tests-1b.R`, the NaN error messages from `reg-tests-2.R`, and the four `\examples` from `seq.Rd`. The rest of the table walks the branches of `seq.default` those do not reach: `from:to` in both directions, all four routes to a single value, a subnormal step, and the overflow rewrite. Expected values are frozen `%.17g` literals generated by `t/seq.R.R`, committed next to the test; the test never calls R. - Python has no equivalent to cross-check against, and that is recorded in the file rather than left as a gap: `numpy.arange` is half-open and carries no fuzz, so `arange(0, 1, 0.1)` is ten values ending at 0.9 where `seq(0, 1, 0.1)` is eleven ending at 1, and `numpy.linspace` is parameterised by length, which is R's `seq(length.out=)` and a different function. - Agreement with R is exact on a `double` perl — the worst relative disagreement over all 378 assertions is 0 on perl-5.44.0 and perl-5.42.3. The other builds cannot be exact, and not because their arithmetic is worse: a long-double or `__float128` perl parses `"0.05"` to its own NV, half an ulp of a double away from the number R read, and one multiply and one add carry that through. The tolerance is measured rather than chosen — 1.11e-16 on perl-5.10.1, 1.50e-16 on perl-5.12.5 and 5.44.0-quadmath, 1.60e-16 on an x87 build — and set at 8e-16, five times the worst of them. - Two cases are asserted as exact integer multiples of a power of two rather than as decimals, because a decimal there would be testing perl's string-to-double instead of `seq`: perl-5.10.1 reads `7.9050503334599447e-323` as `0` and `1e307` five ulp high, while `2**k` is exact on every one of these perls. 0.313 2026-09-01 CDT [`merge`: a `left.on`/`right.on` join died when the right frame reused the key's name] - The result of a join carries one key column, under the left name — R's convention, and not pandas', which keeps both. That left one collision unaccounted for: with `left.on`/`right.on`, a *non-key* column on the right can be named after the *left key*, and it then collides with the output key column rather than with a left data column. The suffix rule only looked at the left frame's data columns, so nothing was renamed and the collision guard fired: merge($parents, $children, 'left.on' => 'name', 'right.on' => 'parent'); # merge: output column 'name' collides; adjust 'suffixes' - That is `tests/reg-tests-1d.R`'s parents/children join, and both references perform it: R suffixes the right-hand copy under `no.dups = TRUE` (the default since R 3.5.0 — "if a `by.x` column name matches one of `y`, the y version gets suffixed as well"), and pandas keeps both key columns so the question never arises. `merge` now suffixes it too, giving R's names exactly: - | | columns | |---|---| | R 4.6.1 | `name`, `sex.x`, `age.x`, `name.y`, `sex.y`, `age.y` | | `merge` before | *croaks* | | `merge` now | `name`, `sex.x`, `age.x`, `name.y`, `sex.y`, `age.y` | - Every input this changes used to croak, so no join that worked before returns anything different. If the suffixes still leave two output columns sharing a name, `merge` still dies rather than hand back a frame with a column missing — which is R's behaviour too (`suffixes = c(".z", ".z")` is an error there). [`merge` is now cross-validated against R's and pandas' own merge suites] - `t/merge.R.pandas.t` (329 tests) takes its cases from the references' test suites rather than inventing them, in the manner of the other `t/*.R.scipy.t` files: - **R 4.6.1** — 28 frames and 83 cases: the examples in `src/library/base/man/merge.Rd` (authors/books, and the `incomparables` example), plus R's own regression cases for `merge` from `reg-tests-1a.R` (PR#1510 `by.x`/`by.y` with multiple matches; the Cartesian product that did not make column names unique in 2.3.0; "merging when NA is a level"; the two character matrices that failed pre-2.0.0; merge on zero-row frames, not allowed ≤ 2.4.0), `reg-tests-1b.R` (the 2.15.0 and 2.15.1 suffixes regressions; the `women` zero-row merges that failed in 2.7.0), `reg-tests-1d.R` (the `by.y` naming case above) and `reg-tests-2.R` (the authors/books joins moved out of `merge.Rd`, and the 2002 case where every column is a join key). - **pandas 2.2.3** — 31 frames and 60 cases from `pandas/tests/reshape/merge/test_merge.py` (`test_intelligently_handle_join_key` GH#733, `test_merge_overlap`, `test_merge_different_column_key_names`, `test_merge_same_order_left_right` GH#35382, `test_left_merge_empty_dataframe`, all ten parametrisations of `test_merge_empty` GH#52777, `test_merge_on_ints_floats`, `test_merge_non_unique_index_many_to_many`, `test_merge_suffix`), `test_merge_cross.py` (all four cross-join tests) and `test_multi.py` (`test_merge_na_keys`, `test_merge_multiple_cols_with_mixed_cols_index` GH#29522). - Each frozen case runs through every input/output shape — AoH, HoA and HoH on the left crossed with the same on the right and with both output shapes, 18 joins per case — and its output column names are checked separately, because a join whose answer has no rows is exactly what several of the reference cases are about and a row-by-row comparison cannot see the names. Beyond the tables the file pins row order against pandas' `sort=False`, R's `by`/`by.x`/`by.y` and pandas' `left_on`/`right_on` spellings, every error path the references document, and double-valued keys with dyadic literals. - The two tables are generated by `t/merge.R.pandas.R` and `t/merge.R.pandas.py`, committed beside the test; the test itself never runs R or python, and needs neither installed. Both blocks reproduce byte-identically from the generators, and the pandas block is byte-identical under pandas 2.2.3 and 3.0.4, so nothing frozen there is specific to a release. - `t/merge.t` also gained the cases a data frame cannot express, and so no reference case can reach: an AoH row missing the key column entirely, a missing non-key cell, the `0` / `'0'` / `'0.0'` key identity, a NUL byte inside a key, a wide character, `Inf`/`-Inf`/`NaN` keys, and blessed frames. [An `undef` join key matches nothing — which is *not* "the pandas `NaN` rule"] - The documentation for `merge` credited its treatment of a missing join key to pandas, and `LikeR.xs` credited it to R's default. Both attributions were wrong, in the same direction: both references match a missing key to a missing key. On `merge.Rd`'s own `incomparables` example the two agree with each other exactly and disagree with `merge`: - | join | R 4.6.1 | pandas 2.2.3 | `merge` | |---|---|---|---| | `by = c("k1","k2")` | 3 rows | 3 rows | 2 rows | | `by = "k1"` | 6 rows | 6 rows | 2 rows | - The behaviour is right and unchanged — it is SQL's rule for a `NULL` key, which is `merge(..., incomparables = NA)` in R, the line `merge.Rd` itself runs, and what `merge`'s documentation describes everywhere else. Only the credit was wrong; R's default is `incomparables = NULL`. The corrected wording says which rule it is and that it is not either reference's default, and `t/merge.R.pandas.t` now asserts the divergence — with R's own `incomparables = NA` answers frozen beside the generalised ones — so changing it later has to be deliberate. [11 leak checks reported Devel::Cover's own counters as leaks] - `t/vif_hoslem.t` failed under `cover` with 37 leaks attributed to a line of `hosmer_lemeshow` that allocates nothing (`next unless defined ... && looks_like_number ...`), and the leaked SVs were bare `IV`s holding values like `98113961175368` — pointers. They are Devel::Cover's per-line counters, which are allocated inside whichever block happens to be running and which `Test::LeakTrace` then counts; the same test reports 0 leaks on a plain perl. - Around ninety test files already guard their leak checks with `$INC{'Devel/Cover.pm'}` for this reason. Eleven did not: `t/age_standardize.t`, `t/dunn_test.t`, `t/effect_sizes.t`, `t/friedman_test.t`, `t/glm_families.t`, `t/ks_test.R.scipy.t`, `t/mcnemar_test.t`, `t/prop_test.t`, `t/tied.frames.t`, `t/vif_hoslem.t` and — for an import it never used — `t/_parse_csv.t`. Four of them failed; the rest passed only by luck, since whether the counters land inside the measured block depends on which lines get their first coverage there, which moves with file order. All of them are guarded now, in the two forms the suite already uses, and the whole suite passes under `Devel::Cover` (143 files, 35008 tests) as well as without it (143 files, 35577 tests — the difference is the leak checks, which still run and still report 0 leaks outside coverage mode). [`pt` near zero and `pf`'s upper tail lost up to eight digits to a cancellation] - `incbeta()`, the regularized incomplete beta every t, F and binomial tail is built on, took only `x` and re-formed `1 - x` by subtraction. Its reflected branch — `I_x(a,b) = 1 - I_{1-x}(b,a)`, taken exactly when `x` is the side near `1` — then needs that complement, and forming it as `1.0 - x` is catastrophic cancellation there. Once `|1 - x|` fell below `2^-53` it collapsed to `0` and took the whole tail with it: - | call | returned | correct to 21 digits | |---|---|---| | `pf(1e-12, 1, 1e6, 'lower.tail' => 0)` | `1` exactly | `0.999999202115638638` | | `pt(1e-6, 1e6)` | `0.5` exactly | `0.500000398942180682` | | `pt(-1e-8, 1)` | `0.5` exactly | `0.499999996816901138` | - Every caller had the complement exactly, and was throwing it away: `pf` builds `x = df1·f/D` and `1-x = df2/D` over one denominator, `pt` has `x = df/(df + t²)` against `t²/(df + t²)`, and `pbinom`'s lower tail is `I_{1-p}(n-k, k+1)` with `p` itself to hand. So the fix is to pass both — `incbeta_xy(a, b, x, y)`, which is R's own split (`bratio()` takes `x` and `y` as separate arguments for this reason), with `incbeta()` kept as the one-argument wrapper for the callers that genuinely have only `x`, such as a bisection midpoint. - Worst error against `mpmath` at `mp.dps = 80`, over a 372-point grid: - | | before | after | |---|---|---| | `pt`, either tail | `3.99e-07` absolute | `1.11e-16` (1 ulp) | | `pf` upper tail | `3.19e-08` relative | `2.74e-11` | | `pf` lower tail | `2.01e-12` relative | `2.46e-13` | - The two tails also add up again: `pf(1e-12, 1, 1e6)`'s lower and upper summed to `1 + 7.98e-07` before, which is what first showed the bug. - This reached `t_test`, whose p-value is `1 -` that tail: a statistic small enough on enough degrees of freedom had its p-value pinned at exactly `1` instead of `1 - 4e-7`. `d_pf` now calls `pf()` rather than repeating its expression, so the two can no longer drift apart on the complement argument. - The one place `incbeta_xy` does not help is the Clopper-Pearson upper bound for a handful of successes in ~1e9 trials, which still carries ~2e-9 of relative error. The cancellation there is in the continued fraction's own argument during bisection, not in a complement a caller could have supplied, so `t/binom_test.R.scipy.t`'s existing note — that fixing it properly means porting `bratio()` — still stands. [`t_test` rejected `conf.level`, the spelling its own documentation lists] t_test(\@x, \@y, 'conf.level' => 0.99); # t_test: unknown argument 'conf.level' - `t_test` accepted only the underscored `conf_level` and `var_equal`, while its parameter table in this file has always documented the argument as `conf.level`, and while every sibling in the module — `var_test`, `wilcox_test`, `prop_test`, `cmh_test`, `glm` and the rest — already took both spellings. `var.equal`, which is what R calls it, was refused too. Both dotted forms now work, and the parameter table records the aliases. [`rank`, `wilcox_test` and `ks_test` are 1.6× to 2.3× faster] - These were still sorting through `qsort()`, whose comparator the compiler cannot inline; the module's own `LIKER_DEFINE_SORT()` introsort and `nv_sort()` were already used by the three internal rankers but not by these. Ordering *n* records costs O(*n* log *n*) comparisons and an indirect call on each is most of what such a sort costs — the figure `LIKER_DEFINE_SORT`'s own comment records is 210µs against `qsort()`'s 355µs on 5000 NVs. - Measured on 20,000 doubles: - | | before | after | speedup | |---|---|---|---| | `wilcox_test` (two samples) | 6.11 ms | 2.64 ms | 2.3× | | `rank` | 2.87 ms | 1.43 ms | 2.0× | | `ks_test` (two samples) | 3.70 ms | 2.37 ms | 1.6× | - `rank_and_count_ties()`, which `wilcox_test`, `ks_test` and five other functions share, had been sorting `RankInfo` records through `cmp_nv3` — a comparator that reads a bare `NV`. That worked only because `val` is the struct's first member and a pointer to a struct is a pointer to its first member: true, but fragile as well as slow. It has a generated ordering of its own now. - What is left in `rank` is no longer the sort. Reading the same 20,000 values back out of an `AV` in plain perl costs 0.37 ms, so the gather loop and the `newSVnv` per result are now most of the call. [Five new cross-validation files, from R's and SciPy's own test suites] - 4,444 tests, taking their cases from the references' suites and documented examples rather than inventing them, in the manner of the existing `t/*.R.scipy.t` files. Expected values are frozen literals; the generators are committed beside each test and are never run by it, so nothing here needs R, python or `mpmath` at install time. - | file | tests | sources | |---|---|---| | `t/t_test.R.scipy.t` | 2437 | `t.test.Rd`; `reg-tests-1a.R:4529` (one group of size one), `reg-tests-2.R:3199`, `reg-tests-1e.R:1985`; SciPy's `TestTTest_1samp`, `TestTTestIndMore`, `TestTTestRel`, `TestTTestCI` | | `t/tukey_aov_prcomp.R.t` | 940 | a 637-point `ptukey`/`qtukey` grid; PlantGrowth and chickwts `TukeyHSD`; mtcars `anova`/`vif`; USArrests `prcomp`; `scale` | | `t/p_adjust.R.t` | 767 | every method in R's own `p.adjust.methods`, on `p.adjust.Rd`'s own p-vector | | `t/friedman_mcnemar_prop_cmh.R.t` | 274 | Hollander & Wolfe (1973) p.140ff; Agresti (1990) p.350; Fleiss (1981) p.139; Agresti's Rabbits and `UCBAdmissions` | | `t/pf_pt_tails.R.mpmath.t` | 26 | `mpmath` at `mp.dps = 80`, plus R on the same grid | - `t_test` had no cross-validation at all before this, which is how the `conf.level` croak above survived. Each documented case is crossed over the whole argument space its function exposes — alternative × `var_equal` × `mu` × `conf.level` × `paired`, `correct`, `exact`, `p` — because a reference case exercised only at its defaults pins one code path out of dozens. - Two things fell out of writing them. - **`t/tukey.t`'s tolerances understate the `ptukey` port by ten orders of magnitude.** It checks `qtukey` to an absolute `1e-3` and the `TukeyHSD` columns to `1e-4`, where the Copenhaver & Holland port actually agrees with R to `3.0e-14` and `2.5e-12` over the new grid. A `1e-3` limit would not notice the port being replaced by a normal approximation, which is the regression it exists to catch. The new file checks it properly; the old one is left alone. - **A tail probability needs an absolute tolerance, not a relative one, across NV widths.** `ptukey`'s internals are deliberately plain `double`, exactly as R's `src/nmath/ptukey.c` has them, so the quadrature does not move with perl's `NV` — but its *argument* `q = |diff| / se` comes from the `aov` mean square, which is computed in `NV`. On the chickwts `horsebean-casein` comparison, where `p adj` is `3.07e-08`: - | NV width | `p adj` | relative | absolute | |---|---|---|---| | `double` | `3.0701967967949884e-08` | — | — | | x87 `long double` | `3.0701966635682254e-08` | `4.34e-08` | `1.3e-16` | | `long double` | `3.0701956643675032e-08` | `3.69e-07` | `1.1e-14` | | `__float128` | `3.0701956643675032e-08` | `3.69e-07` | `1.1e-14` | - Every width agrees to about `1e-14` absolute, which is all a probability in the `1e-8` tail can be asked for. Conversely `pf` and `pt` come out three to four orders *more* accurate on the wider widths (`2.74e-11` → `1.1e-15` for `pf`'s upper tail), which is the evidence that what is left there is the continued fraction's convergence and not another cancellation — a cancellation does not improve with the working precision. [`qf` is more accurate than R's own `qf` in the far lower tail] - Writing the file above turned up a disagreement in the other direction, so it is recorded rather than quietly reconciled. Asked for the F quantile deep in the lower tail, R bottoms out at a fixed resolution: - | | `qf(1e-8, 1, 2)` | |---|---| | `mpmath`, `mp.dps = 80` | `2.0000000000000003e-16` | | `qf` | `2.0000000000000000e-16` | | R 4.6.1 `qf` | `4.4408920985006262e-16` — which is `2^-51` | - R is out by 122% there. Over the 108-point grid in `t/pf_pt_tails.R.mpmath.t`, `qf` is within `7.6e-15` relative of 80-digit truth at every point and R is out by as much as `1.22`, on 17 of them. The test asserts `qf` against `mpmath` and separately asserts that R really is the worse of the two, so that "fixing" `qf` towards R would fail loudly instead of passing quietly. - Also recorded there: `qtukey` is *not* an exact inverse of `ptukey`, in R or here. R's `qtukey.c` is a secant iteration that stops once successive iterates differ by less than a hardcoded `const static double eps = 0.0001` — an absolute `1e-4` in `q`, not a relative tolerance on `p` — so `ptukey(qtukey(p))` recovers `p` only to about `1e-7`. R's own worst round trip over that grid is `1.2744607e-07`, and this port's is the same to the digits printed, which is the strongest evidence in the file that the port really is faithful rather than merely close. 0.312 2026-08-28 CDT [80-bit x87 arithmetic broke four test files on 32-bit x86] - A CPAN smoker on `i686-linux-64int` (perl 5.32.1) failed 73 subtests across `t/bedroc.python.t`, `t/quantile.R.t`, `t/var_sd_cov.R.t` and `t/wilcox_test.R.scipy.t`. None of it reproduced here: the local matrix varies how wide an NV is *stored* — double, long double, `__float128` — but every perl in it is x86-64, so all five agree on doing the arithmetic in SSE at exactly that width. - 32-bit x86 does not. It computes doubles in the x87 register file at 80 bits and rounds only when a value is spilled to memory, so `FLT_EVAL_METHOD` is 2 where an SSE build reports 0. C99 says a cast or an assignment rounds to the declared type, but gcc implements that only under `-fexcess-precision=standard`, and `-std=gnu99` — which the C99 probe in `Makefile.PL` tries first — selects `-fexcess-precision=fast`, which keeps the extra bits. A product therefore carried more precision than a double has, and every place that split one into an integer part and a remainder picked a different integer: - | expression | SSE | x87 | |---|---|---| | `ceil(0.1 * 100)` | `10` | `11` | | `ceil(0.05 * 60)` | `3` | `4` | | `1000 * 0.007` | exactly `7` | `7.000000000000000146` | - One cause, four faces. `bedroc` and `auroc` size a bucket with `ceil(frac * N)`, so every count came back one too high — each wrong enrichment factor in the smoker's report has an eleven under it where a ten belongs, and `n.active` was 4 where 3 was wanted, 21 where 20 was. `quantile` forms `h = (n - 1) * p` and interpolates whenever `h` is not an integer, so at *n* = 1001 and *p* = 7/1000 it interpolated between two order statistics where R returns one of them exactly. `var` and `sd` disagreed with themselves in the last digit between the tied-array and plain-array paths. And `wilcox_test` found its pseudomedian by a root search over a step function, where a perturbed iterate does not move the answer by an ulp but lands on a different flat step: - | case | before | R 4.6.1 | |---|---|---| | `b12`/`c12` paired, `mu = -1` | `-0.64836734615018388` | `-0.70304928468751926` | | `a4`/`a9`, `mu = 0.5` | `0.43960806666603025` | `0.4690926266724697` | - The fix is in two layers. `Makefile.PL` and `dist.ini` now trial-compile `-fexcess-precision=standard` and add it when the compiler takes it, alongside the existing per-vendor C99 probe; gcc and clang both exit non-zero on an `-f` switch they do not implement, and MSVC is skipped as before, needing nothing — it has defaulted to SSE2 on x86 since VS2012, and `/fp:precise` rounds to the declared type. That flag on its own fixes all 73. It is still not the whole answer, because it is not everywhere: an older clang, a vendor cc, or any compiler whose only C99 mode is a wide-register one will not take it, and this module is expected to build on all of them. - So `LikeR.xs` also gains `nv_narrow()`, which forces the store through a `volatile NV` by hand at the seven places where a product is split into an integer part and a remainder: both passes of `quantile` — which must agree, or the partial sort places the wrong order statistics — plus `dens_quantile7`, `_qcut_core`, and the bucket counts in `bedroc` twice and `auroc` once. Those are the sites where the extra bits change *an answer* rather than its last digit, and by itself it fixes 63 of the 73. The other 10 are last-digit work — `var` agreeing with itself between the tied-array and plain-array paths, and the root search landing on the same step — which is what the flag is for and what no reasonable amount of hand-placed rounding would buy. - Rounding to the NV's own width is the right answer at every width, not a concession to double: 7/1000 times 1000 rounds to exactly 7 in a double and in a long double alike. So both layers are inert on the long-double and quadmath builds, which is what the matrix confirms. - No new tests were written, because the four cross-validation files that caught this are already the regression test — they are pinned to R and SciPy and they fail without the fix. What was missing was a build that could run them the way the smoker did, so `./test.all.perls.pl` gains a first-class `+x87` row that rebuilds the newest plain perl with `-mfpmath=387`. It reproduced 63 of the 73 subtests by test number. One row is enough and it is skipped where there is nothing to catch — an x87 register is exactly as wide as a long-double NV, and `__float128` never touches the x87 stack — and `--no-x87` turns it off. - That row cannot see the other half of 32-bit x86, and it is worth naming so the gap is not mistaken for coverage: i386 returns a double in `st(0)`, so a callee hands its caller all 80 bits, where the x86-64 ABI returns in `xmm0` and narrows for free. That is the half that moved `wilcox_test`, which is why it was the one failure that could not be reproduced here at all; the assembly shows plain `-std=gnu99` emitting `fdivl` and `ret` with no rounding between them, and the flag inserting the `fstpl`/`fldl` round-trip that narrows the result. Only a real 32-bit perl will exercise it. [`cor_test()`: five disagreements with R, one of which moves p-values] - The tie-corrected variance is the one that matters. `cor_test(method => 'kendall')` divided the Kendall score by `n(n-1)(2n+5)/18`, which is R's `var_S` only when there are no ties — and with no ties and *n* < 50 the branch is not reached at all, so the correction was missing in exactly the case that needs it. R builds it from the tie-group sizes of each vector: var_S = (v0 - vt - vu)/18 + v1/(2n(n-1)) + v2/(9n(n-1)(n-2)) - On fifteen points of tied integer data: - | | 0.311 | R 4.6.1 | |---|---|---| | z | `-1.8805123053604953` | `-2.0721033457107345` | | p (two-sided) | `0.060038290909579115` | `0.038255804392841472` | - A factor of 1.57 on the p-value, across 0.05. Integer scores, counts and Likert data are all ties, which is most of what a Kendall correlation is asked for. - Four more, in the same function: - **Ties of odd group size were not ties.** The Spearman branch decided `TIES` by looking for a fractional average rank, and a group of three averages to a whole number: `y = (1,1,1,2,2,2,3,3,3,4,4,4)` ranks as 2, 2, 2, 5, 5, 5, 8, 8, 8, 11, 11, 11 and was called tie-free, which sent it to the exact branch R reserves for data that has none. It is now R's test — a repeated *value* — taken from the tie walk `rank_data()` was already doing. - **Spearman's `statistic` under ties.** R always forms *S* as `(n³ − n)(1 − ρ)/6`; the sum of squared rank differences is the same number only without ties. 0.311 reported the sum: `793.5` where R gives `865.39724699477085`. It is now the identity, on every path. - **The Pearson confidence interval ignored `alternative`.** R switches on it — `less` is `(-1, tanh(z + σ·qnorm(cl))]`, `greater` is `[tanh(z - σ·qnorm(cl)), 1)`. The two-sided interval came back whatever was asked for: `[0.5217431448512, 0.9680507713838]` for `alternative => 'greater'` where R gives `[0.6029901323843, 1]`. The p-value was never wrong, only the interval, which is the kind of error that reads as an answer. There is also no interval below *n* = 4 now, as in R. - **The continuity correction at a score of zero.** R applies `S <- sign(S) * (abs(S) - 1)`, and `sign(0)` is 0, so a score of exactly 0 stays 0. Subtracting a signum that treats 0 as negative moved it to +1: thirty points whose score is 0 reported `z = 0.019410388389502`, `p = 0.98451372323408` where the answer is `z = 0`, `p = 1`. - One divergence went the other way and is now gone. The exact Kendall distribution summed its lower tail from `dp[max_inv]` downwards, the end of the array that has accumulated every rounding the recurrence made; the distribution of inversions is symmetric, so both tails are now summed from `dp[0]`. Perfect discordance at *n* = 10 returned `2.7557319223626736e-07` and is now exactly 1/10!, `2.7557319223985891e-07`, which is R's value. In the mirror case R is the one that is 9.3e-11 out and `cor_test` gives the exact double; `t/cor_test.kendall.R.t` records that and asserts the combinatorics. - `t/cor_test.kendall.R.t` (1236 assertions) and `t/cor_test.pearson.ci.R.t` (274) are new and generated from R; `t/cor_test.spearman.R.t` gained the tied rows it had been quoting as a recorded divergence. [Kendall's counts, in O(*n* log *n*)] - `cor_test(method => 'kendall')` walked every (*i*, *j*) pair to count concordances — the O(*n*²) loop `cor()` had already replaced with Knight's algorithm in 0.31. Both now come from one routine, which returns the tie moments R's `var_S` needs along with the pair counts, so the correctness fix above and this are the same edit: - | *n* | `cor_test` before | `cor_test` now | |---|---|---| | 4 000 | 0.059 s | 0.0007 s | | 16 000 | 0.93 s | 0.0031 s | | 64 000 | 14.73 s | 0.0139 s | - `cov(..., 'kendall')` is the same sum over the full *n* × *n* space, so it is 2(C − D) and comes from the same counts: 1.85 s to 0.0061 s at *n* = 32 000. R's own `cov(method = "kendall")` is O(*n*²), so this is now faster than the reference rather than merely equal to it. [`hist()`: R's `pretty()` breakpoints] - R does not cut the range into `breaks` equal pieces. It treats that number as a suggestion and asks `pretty()` for round numbers, so the axis ends outside the data and the interval count need not be the one requested: hist([1,2,2,3,4,7,9,10,11,15]) breaks 0 2 4 6 8 10 12 14 16 was 1 5.75 10.5 15.25 20 mids 1 3 5 7 9 11 13 15 was 3.375 8.125 12.875 17.625 - `R_pretty()` is ported from `src/appl/pretty.c` with `pretty.default()`'s parameters and `hist.default()`'s `min.n = 1`; the counts come from a binary search over the breakpoints carrying R's 1e-7 fuzz, which is what puts a value sitting a rounding error above a break into the bin below it. The default bin count is R's `nclass.Sturges`, ⌈log₂ *n* + 1⌉ — the ceiling is part of it, and a truncating cast had been turning `log₂(13) + 1 = 4.70` into 4. - `t/hist.R.t` pins all four returned vectors across thirteen datasets — a single point, a constant column, data spanning zero, data of order 1e-3 and 1e6 — crossed with eight `breaks` settings. Every value matches R exactly, to the last bit. [Arguments that are not numbers] - `SvNV()` turns a string that is not a number into 0 without a word, because the `numeric` warning it would raise belongs to the caller's scope and an XSUB does not run in one. So `min(1, 2, 'abc')` returned 0 — a value that is not among its arguments and is not the smallest of them — and `mean(1, 2, 'abc')` returned 1. `undef` had always been refused; this was the hole beside it. - `mean`, `sum`, `min`, `max`, `median`, `var`, `sd`, `skew`, `kurtosis`, `quantile`, `scale` and `hist` now croak, naming the argument or the cell. An object that overloads its numeric conversion still counts as a number; a reference that overloads nothing is now refused too, where `SvNV()` used to fold in its address. The cross-validation file that had been asserting the old behaviour — `t/read_table.na_strings.t`, on an unmapped literal `NA` — asserts the croak instead, which is a better signal than the warning it could not see. [`NaN` in `median()` and `quantile()`] - No comparison sort can place a `NaN`: every comparison against one is false. So which value ended up in the middle depended on where in the array it sat. `median()` of 1..50 with one `NaN` in it returned 25.5, `median([5, 1, NaN])` returned 5, and `median([1, 2, NaN, 4])` returned `NaN`. `quantile()` of that same 50 answered `13.25` for the 25% and `NaN` for the 75%. - Both now propagate — one `NaN` in, `NaN` out — which is what `sum`, `mean`, `var`, `sd`, `min` and `max` already did here, and what keeps the 50% quantile equal to `median()`. `undef` is still dropped by `quantile()`, as documented. - The five bandwidth selectors disagreed with each other about the same input: `bw_ucv`, `bw_bcv` and `bw_sj` rejected a non-finite value a layer down, while `bw_nrd0` and `bw_nrd` returned a number — 7.5250570111922 for 1..50 with two `NaN`s, computed from a variance and an interquartile range that are both `NaN` via a `min()` that a `NaN` silently loses. All five now refuse it in the shared reader, which is also what keeps a `NaN` out of the `qsort()` there: `cmp_nv3()` cannot order one, and an inconsistent comparator makes `qsort()` undefined behaviour rather than merely inaccurate. [Ragged frames] - A HoA frame whose columns have different lengths was fitted on whichever column `hv_iternext()` reached first, so how many rows `lm()`, `glm()` and `prcomp()` used moved with hash order from run to run: `{ y => [1..6], x => [1,2,3] }` returned a perfect fit on three rows with `df.residual = 1` and said nothing. All three now croak, naming the column and both lengths, as `csort()` already did; `prcomp()` checks a ragged AoA too. - This is stricter than R on purpose. `data.frame()` recycles a short column when its length divides the longest — `data.frame(y = 1:6, x = 1:3)` quietly fits on `x` repeated twice, and only `1:4` is an error there. Inventing observations is not something to do quietly in a model fit. [`scale()`: named arguments] - Every other function in this module takes its options as trailing name/value pairs. `scale()` read only a trailing hash reference, so a flat pair fell through into the data, where the option name numified to 0 and was scaled along with everything else: `scale([1,2,3], center => 0)` returned five numbers for a three-element input. Both forms are now accepted and agree; an unrecognised name is an error rather than a data point. [`survfit()`: where the curve reaches zero] - Greenwood's next variance term is `d / (n.risk × (n.risk − d))`, and the last event takes every remaining subject, so `n.risk == d` and the variance of *S*(*t*) is not defined from there on. R reports `std.err` as `Inf` and both confidence limits as `NA`. All three came back 0, which is a number where there is no answer, and a plausible-looking one. They are `undef` now. [Two smoker failures, both in the tests rather than the module] - `LikeR.xs` is unchanged here. Two failures came back from CPAN smokers on 0.31, and in each case the assertion had been pinned to something that is a property of the perl running it rather than of `Stats::LikeR` — and this machine's build happens to satisfy it, which is why neither showed up before the reports arrived. - `t/avals.t` was counting `FETCH`es on a tied cell: - The tied-cell block built its frame by copying a tied scalar into an anonymous hash, and used a `FETCH` that counted up so that a second fetch would be visible: sub FETCH { my $s = shift; return $$s++ } # a different value each FETCH my $tied; tie $tied, 't::AlwaysSeven'; my @got = avals([ { v => $tied } ], 'v'); is($got[0], 7, 'a tied cell is fetched once, at extraction time'); - `{ v => $tied }` fetches the cell while the frame is being built, so what `avals()` copied was already a plain number: the test never exercised the thing its own name described, and what it actually measured was how many times that perl runs get-magic on one `sv_setsv()`. That count is not fixed. Measured directly, `perl-5.10.1`, `perl-5.12.5` and `perl-5.44.0` each run it once and the caller sees 7; a smoker on 5.18.4 (`x86_64-linux-thread-multi`) reported 8, i.e. two runs on the one copy, and a second report of the same two subtests followed. - The cell is now tied in place, so nothing touches it before `avals()` does, and `FETCH` returns a constant, so nothing depends on how many times it is called: my %row; tie $row{v}, 't::TiedCell', 7; - What the block asserts is now what it always claimed: building the frame fetches nothing, extraction fetches at least once, the fetched value rather than the tied SV comes back, the copy carries no tie magic, re-reading it does not fetch again, and it is exactly what `vals()` returns for the same frame. The HoH branch and the HoA branch — which fills `AvARRAY` directly and so has its own `newSVsv()` — get the same treatment, 44 subtests to 55. - `t/01.t` compared a number of order 50 against an absolute `1e-14`: - The other report was the `lm()` F statistic, exactly 54 by algebra: # got: 54 # expected: 54; diff = 1.4210854715202e-14 - `is_approx()` treated its `$epsilon` as an absolute bound, and 54 sits where the spacing of a double is `ulp(54) = 2**-47 = 7.105e-15`. A bound of `1e-14` is therefore 1.41 ulps — the call site was demanding bit-exactness of a number of order 50, and the smoker's build simply reassociated the sum and landed two ulps out (`1.4210854715202e-14` is exactly `2**-46`). The same `1e-14` written against a p-value of order `1e-9` is some 10⁵ ulps, which is what the number was chosen to mean. - The allowance is now scaled by the size of the expected value, floored at the absolute bound so that comparisons against 0 — and against anything below 1 — keep exactly the tolerance they had: my $scale = abs($expected) > 1 ? abs($expected) : 1; my $tol = $epsilon * $scale; - The F statistic gets 76 ulps of headroom in place of one. The failure `diag()` also prints the tolerance now, so the next report of this shape can be read without opening the file. - Every `is_approx`-style call in all 138 test files was then parsed and its tolerance compared against `ulp(expected)`. Nine sites sat under eight ulps, all of them in `t/01.t` and all now covered by the scaling: - | `t/01.t` line | expected | tolerance | was | |---|---|---|---| | 1149, 1370, 3056, 3081, 3100, 3399 | 54, 37, 58, 58, 40, 36 | `1e-14` | 1.41 ulp | | 2216 | 31 | `1e-14` | 2.81 ulp | | 3519, 3561 | 177.504798464491358 | `1e-13` | 3.52 ulp | - Call sites with a tolerance of `1e-13` or tighter against a computed value exist in `t/prcomp.t`, `t/t_test.t`, `t/value_counts.t` and `t/wilcox_test.t`; each was checked by hand and compares a p-value, a count, an integer rank sum, or an `sdev` of about 2.8, so those bounds are thousands of ulps wide and are left alone. - Both files pass on `perl-5.44.0`, `perl-5.10.1` and `perl-5.12.5` — `t/01.t` 1399 of 1399, `t/avals.t` 55 of 55 — the two older perls built in private trees so that the distribution root's `Makefile` and `blib` stay as they were. 0.311 2026-08-28 [avals] - new function to make code cleaner. `avals` is the same as `vals` but returns an array instead of an array reference, so instead of `@{ vals($df, 'colname') }` it will just be `avals($df, 'colname')` to prevent the need for bracketing [Correctness and portability fixes] - Nine defects found by cross-validating against R 4.6.1 and mpmath, by running the suite under UBSan and AddressSanitizer, and by reading for the conventions in `CLAUDE.md`. Each is pinned by a test that fails without the fix: `t/cor_test.spearman.R.t`, `t/pt_qt.tails.R.t`, `t/minmax.nan.R.t`, `t/var_test.ci.R.t` and `t/portability.bugs.t`. - `cor_test(method => 'spearman')` could not return: - `exact => 1` ran an unbounded permutation walk. R has never enumerated past *n* = 9 — `n_small` in `src/library/stats/src/prho.c` — whatever `exact` says, but here the caller's flag was taken literally, and *n*! is 6.2 × 10²³ by *n* = 24: - | *n* | before | now | |---|---|---| | 12 | 5 s | 0.0 s | | 13 | > 1 min | 0.0 s | | 24 | never returned | 0.0 s | - The enumeration is now capped at *n* ≤ 9 and 10 ≤ *n* ≤ 1290 goes to the AS 89 Edgeworth series, which is what R's `prho()` does. Past 1290 — R's own bound, where `n^3 - n` overflows an `int` — it falls back to the asymptotic *t*, as before. - Spearman p-values were R's `exact = FALSE` for every *n* ≥ 10: - `exact` defaulted to `(n < 10)`. R's default is `TRUE`, and R serves it from `prho()` for every *n* up to 1290, so the old default silently took R's approximation branch: - | *n* | before | R 4.6.1 | |---|---|---| | 32 | `1.7639963277557958e-07` | `1.1513643381176551e-06` | | 16 | `5.0655870675677e-17` | `5.7949654941386787e-06` | - Across 45 randomly generated pairs the p-value now agrees with R to `3.5e-14`. - Spearman's `statistic` changed meaning at *n* = 10: - It was *S* from the exact branch and the *t* statistic from the other. It is now *S* on every path, as R reports it. This is a change to a documented field — the note saying to compare `estimate` and `p.value` rather than `statistic` described the old behaviour and is gone. - `qt` saturated, and `pt` collapsed to zero, in the far tail: - `pt_upper()` computed the tail as `I_z(df/2, 1/2)` with `z = df/(df + t²)` and nothing else. `z` underflows well before `t²` overflows, so the tail became a flat `0`; and because `qt_tail()` bisects against it, and its bracket stopped at `sqrt(DBL_MAX)`, the quantile converged onto the ceiling and returned it — a plausible number, not an error: - | call | before | R 4.6.1 | |---|---|---| | `pt(-1e160, 1)` | `0` | `3.1830988618378451e-161` | | `pt(-1e300, 1)` | `0` | `3.1830988618377598e-301` | | `qt(1e-160, 1)` | `-1.3407807929942597e154` | `-3.1830988618379068e159` | | `qt(1e-300, 1)` | `-1.3407807929942597e154` | `-3.183098861837907e299` | | `qt(1e-16, 0.1)` | `-1.3407807929942597e154` | `-1.6044257056664863e156` | - `pt_upper()` now carries the Abramowitz & Stegun 26.5.4 branch from R's `nmath/pt.c`, which evaluates the tail in logs once `1 + (t/df)·t` passes `1e100` and so has nothing left to underflow, and the bracket runs to `NV_MAX/2` and reports `Inf` when the quantile genuinely exceeds what an `NV` can name. `pt` now matches R exactly from `-1e150` to `-1e300`. - `min` and `max` skipped a `NaN` unless it came first: - Every comparison against a `NaN` is false, so `if (v < mn)` passed over one and kept a real number — but a `NaN` in the first position stayed. The answer depended on where in the array it sat: min([NaN, 1, 2]) was NaN min([1, 2, NaN]) was 1 min([1, NaN, 2]) was 1 max([1, 2, NaN]) was 2 - R gives `NaN` for all of them, and `sum`, `mean`, `var`, `sd` and `median` here already did. Both now propagate `NaN` from any position, in the flat argument list and the array-reference forms alike. - `var_test`'s confidence interval lost up to `2.5e-11`: - It divided by a private `qf_bisection()` that stopped on an absolute `high - low < 1e-12`. The F quantile it divides by is not of order 1 — at `conf.level = 1 - 2^-20` on nine and nine degrees of freedom it is about `0.016` — so an absolute stop is a relative error that grows as the quantile shrinks: - | `conf.level` | before | R 4.6.1 | |---|---|---| | `1 - 2^-20` | `[16085387951.278957, 62168223924585.172]` | `[16085387950.871651, 62168223923448.094]` | - It now uses the same `d_qf()` that the exported `qf` does, which has a relative stop and agrees with mpmath at `mp.dps = 60` to about `1e-16`. Over 864 generated cases the interval agrees with R to `2.1e-14`. `qf_bisection` had no other caller and is gone; carrying two F quantiles of different accuracy in one file was the underlying defect. - `var_test`'s p-value was formed by subtraction: - It computed the upper tail as `1 - pf(F)`, which is what [F and z tail p-values](#f-and-z-tail-p-values) says every F test here stopped doing — `var_test` was listed neither among those converted nor among the three known to remain subtractive. It now uses `pf_upper()`, which builds its own complement as `df2 / (df1·F + df2)` without subtracting. R writes the subtraction too, so in the far tail this is now more accurate than R; measured against mpmath over the same 864 cases: - | | worst relative error | |---|---| | `Stats::LikeR` | `1.854e-14` | | R 4.6.1 | `1.0` | - R's worst is `1.0` because `1 - pf(F)` reaches a flat `0` once the lower tail is within an ulp of `1`. On `c(3,1,4,1,5,9,2,6,5,3,5)` against `c(1,2,4,8,16,32,64)/1024`, two-sided: mpmath `1.08732387158297135e-11`, `Stats::LikeR` `1.0873238715829779e-11`, R `1.0873080213968933e-11`. - `%zu` would have failed on MSVC, in two places silently: - Nine plain `snprintf` calls used the C99 `%zu` length modifier, which MSVC's older CRT does not implement — it prints the literal text. Two of the nine build hash keys, not messages: `csort`'s AoA → AoH and AoA → HoA conversions name their columns `"0" .. "ncols-1"`, so on such a build every column would have been stored under the key `"zu"` and all but one lost. A third is the `Index N` group label that `oneway_test` returns to the caller; the rest are counts inside `croak` text. All 9 now use `my_snprintf` with `UVuf`, the modifier Configure picks for the build's own `UV`. - `ssize_t` does not exist outside POSIX: - Seven declarations used the lowercase POSIX spelling. Perl never defines it — there is no `typedef` or `#define` for it anywhere in the perl source, and `win32/win32.h` and `config_H.vc` do not supply one either — so the file compiled only because glibc's `` happens to. An MSVC build failed at the first one. They are `SSize_t` now, or `size_t` where the value is a length that cannot be negative, matching the 430 declarations that already spelled it that way. - Two zero-length calls on null pointers: - `qsort` and `memcmp` are both declared to take non-null pointers whatever the element count, so passing `NULL` is undefined even for zero elements: UBSan reports it, and a compiler may infer non-nullness and delete a later check. `p_adjust` reached `qsort(NULL, 0, …)` whenever no row hash had any keys, and `drop_duplicates` reached `memcmp(NULL, NULL, 0)` when every key seen so far had been zero-length. Both were reachable from one line of Perl — `p_adjust([{}, {}])` and `drop_duplicates([{}, {}])`. The suite is now clean under both UBSan and AddressSanitizer. - `malloc`/`free` in the `write_table` announcement: - The confirmation line allocated with libc's `malloc` when the path pushed it past a 512-byte stack buffer — the only two such calls in the file, where everything else uses `Newx`/`Safefree` so that the module uses whatever allocator its perl was built with, which on a `-Dusemymalloc` perl is not libc's. `Newx` is not the substitute here, since it croaks on failure and dying inside the confirmation line after the table has been written is worse than splitting the line; the long-path case now takes the same three-write path that the old out-of-memory branch did. [read_table] - `na.strings` / `na_values`: - `read_table` can now be told which field texts mean "missing". It is one option under three names — R's `na.strings`, pandas' `na_values`, and `write_table`'s own `undef.val` — so it can be spelled whichever way the surrounding code already does. Each takes a string or an array reference, each means the same thing, and passing more than one is an error the way `sep` and `delim` together are: read_table('cohort.csv', 'na.strings' => 'NA'); read_table('cohort.csv', na_values => ['NA', 'N/A', 'NULL', '-']); read_table('cohort.csv', 'undef.val' => 'NA'); - An empty field has always read as `undef`. Nothing else did, and there was no way to ask for it, so a file could not be read back the way it had been written: `write_table` renders an `undef` cell with `undef.val`, which 0.302 listed as a known limitation whose closing sentence proposed exactly this option. The round trip now closes, and `undef.val` is accepted as the third name so that both halves of it can use one spelling: write_table($rows, 'out.csv', 'undef.val' => 'NA'); my $back = read_table('out.csv', 'undef.val' => 'NA'); # undef again - `write_table`'s `undef.val` is a single token, being what it writes; `read_table`'s takes a list, being every token it should recognise. - The consequence had been quiet rather than loud. An unmapped `NA` is an ordinary string, and Perl numifies a string to `0`, so it did not stop `mean` or `sd` — it pulled them toward zero, warning once per cell on `STDERR` and dying only under `warnings FATAL => 'all'`. On the `titanic.csv` in the repository root, whose 2207 rows carry 3690 literal `NA` cells across six columns: - | column | `NA` cells | mean, unmapped | mean, mapped | |---|---|---|---| | `age` | 2 | 30.416855 | 30.444444 | | `fare` | 916 | 19.540347 | 33.404760 | | `sibsp` | 900 | 0.295877 | 0.499617 | | `parch` | 900 | 0.228364 | 0.385616 | - The mapped column is `mean(d[[col]], na.rm=TRUE)` in R 4.6.1 to every digit shown, for all four. - Three choices are worth stating, because each follows R where pandas differs, and each is asserted in the test rather than left to be discovered: - **It is off by default.** R's `read.table` defaults to `na.strings = "NA"` and pandas recognises a 19-token set, but `read_table` recognises nothing beyond the empty field unless asked. A literal `NA` is real data in every script written against an earlier release, and `t/read_table.t` pins that; changing the default is a separate decision from adding the option. - **The list replaces the set rather than extending one.** `'na.strings' => 'baz'` maps `baz` and leaves `NA` and `NaN` alone. pandas adds `baz` to its defaults and maps all three. - **The match is on the exact field text**, case-sensitively and with no whitespace stripped. `' NA'` is not `'NA'`, and `'-999.000'` is not `'-999.0'` — pandas numifies both sides and matches them; R compares the string and does not, and neither does this. - The mapping happens where the empty-field rule already did, so it reaches every `output.type`, the `.xlsx` reader, and a value a `filter` writes back. The header is never mapped, so a column may still be named `NA`. A `filter` runs after the mapping and so sees `undef` rather than the token, which makes `sub { defined $_ }` the way to keep the rows that have a value. For `hoh`, a mapped row-name cell is a missing row name and is refused exactly as an empty one is. - Everything above is checked in the new `t/read_table.na_strings.t` (95 tests), whose cases are taken from the reference suites rather than invented here: R's `tests/reg-tests-1a.R` — the `read.table(na.strings="foo")` case at 1911-1918 and PR#6781 at 2899-2901 — and `src/library/utils/man/read.table.Rd`, plus pandas 2.2.3's `pandas/tests/io/parser/test_na_values.py` (`test_string_nas`, `test_detect_string_na`, `test_non_string_na_values` for gh-3611, `test_default_na_values`, `test_custom_na_values`, `test_na_trailing_columns`, `test_na_values_scalar` for gh-12224, and `test_na_values_dict_aliasing`). The expected values are frozen literals with their provenance in the file header, so the test needs no R and no Python to run. Three divergences from pandas are asserted deliberately, the three above; a fourth is that a row short of the header is still an alignment error naming the row rather than the NA padding pandas does. - The full suite is 130 files and 26,444 tests, and `./test.all.perls.pl` passes on all five local perls — `5.10.1`, `5.12.5` (long double), `5.42.3`, `5.44.0` and `5.44.0-quadmath` — with no warnings on any of them. [Distribution functions] - Eight new functions — `qnorm`, `pt`, `qt`, `pchisq`, `qchisq`, `pf`, `qf` and `pbinom` — with R's names, R's positional-or-named parameters and R's `lower.tail` / `log.p` flags. Before this, `pnorm` and `dnorm` were the whole family, so a Wald bound or a likelihood-ratio test could not be finished from outside the module: the critical value simply was not expressible. See [Distribution functions](#distribution-functions) for the details; two are worth pulling out here. - No tail is formed as `1 -` the other one. That subtraction discards everything below machine epsilon, which is the range a p-value is interesting in, so each function is routed to the parameterisation that computes the requested side directly — `pt(q)` is `pt_upper(-q)` by symmetry, and `pchisq` lower goes through a new `igam()` (the lower incomplete gamma that `igamc()` had always formed internally and thrown away): pchisq(1e-30, 1); # 7.978845608028654e-16, where 1 - igamc gives 0 - `qf` disagrees with R in the far tail of *F*, and is right. Arbitrated by `t/distributions.mpmath.py` at `mp.dps = 60`, bisecting the defining equation rather than calling a library inverse: - | | R 4.6.1 | this module | |---|---|---| | `qf(2^-20, 1, 10)` | 9.9e-4 out | 1.0e-15 out | | `qf(2^-20, 1, 1)` | 5.2e-5 out | 4.1e-15 out | | `qf(2^-16, 1, 10)` | 2.7e-6 out | 6.9e-16 out | - R's own `pf` settles it: `pf` of this module's answer returns the target probability to 1.6e-15, and `pf` of R's own `qf` answer returns it to 4.9e-4. It is R's inverse in that tail, not its CDF. Thirteen such rows are asserted twice in `t/distributions.R.scipy.t` — right against the 60-digit value, and still disagreeing with R — so if a future R fixes `qf`, or this module regresses onto R's answer, the row fails rather than quietly agreeing. Three rows go the other way and are recorded the same way rather than tolerated. - `log => 1` on a `p*` function is `log()` of the probability already computed, not a log carried through the series, so a tail that rounded to `1` logs to `0` where R reports a tiny negative number. Four of SciPy's six `chi2` log references land in that case and are pinned as exactly `0` with the true value named. `ncp` and `qbinom` are not implemented. - `t/distributions.R.scipy.t` is 2041 tests: a frozen table of R 4.6.1 values at dyadic inputs (generated by the committed `t/distributions.R.scipy.R`), SciPy 1.18.0's own mpmath cases from `TestChi2` and `TestStudentT`, and R's own round-trip identities from `tests/d-p-q-r-tests.R`. Every tolerance is the worst error actually measured, printed by the test's own diagnostic. - Ten figures now sit with the documentation — one per function, `dnorm` and `pnorm` in their own sections and the eight new ones under [What each one does, in one picture](#what-each-one-does-in-one-picture). Each shades the region its function integrates and writes the integral it evaluates, so the `p*`/`q*` pair reads as one identity answered from opposite ends: given the boundary, return the area; given the area, return the boundary. `pbinom` is drawn as bars with a summation sign rather than an integral, because a binomial is discrete and drawing an integral over a histogram would misstate what it computes. `distribution.plots.pl` regenerates all ten; like the other plot scripts it is author-only and `[PruneFiles]` keeps it out of a release. - Two portability problems surfaced while pinning this and are worth recording, because neither is about the C code. Perl 5.10.1's string-to-number conversion does not parse a 17-digit literal exactly — it reads `0.9999847412109375` one ulp out, and reflecting a probability onto `1 - p` turns that ulp into 7.3e-12 relative — and it reads `9.332636185032188789901e-302` as `8.999999999999996e-302`, a 3.7% error. So every argument and every exactly-representable expected value in `t/distributions.R.scipy.t` is written as a dyadic ratio, `M/2^K`, which rebuilds by integer division on any perl. And where R's frozen value is `0` because a double underflowed, a long-double or `__float128` build computes the real number instead — `pbinom(0, 1000, 0.75)` really is `0.25^1000` = 8.7e-603 — so the test accepts that as agreement with a reference that could not go lower, rather than failing the wider build for being right. The same applies to `log => 1` near a probability of 1, where the resolvable accuracy is `NV_EPSILON / |log p|` and the tolerance is scaled off measured machine epsilon accordingly. [Every confidence limit now comes from the same critical value] - Adding `qnorm` made a discrepancy visible that had always been there. Nine places in `LikeR.xs` built their own `z` by calling `inverse_normal_cdf()` — Moro's rational approximation, of which the file's own comment said: *"good to about 1e-9, which is fine for a confidence limit"* — while `qnorm` went through the Newton-polished `normal_quantile_hp()`. Which function you asked therefore decided how many digits you got, and a Wald bound written by hand from `qnorm` differed from the one `glm` reported in the tenth digit. - It is one function now. `std_qnorm()` seeds with Moro, polishes with Newton against the `erfc`-based normal CDF, and reflects a probability above `0.5` onto `1 - p` before inverting; `glm`, `cor_test`, `prop_test`, `epi_2x2`, `cmh_test`, `roc`, `survfit`, `coxph` and `wilcox_test` all call it, `qnorm` *is* it, and two near-duplicates are gone: `wilcox_qnorm()` in the XS, which was the same Newton loop without the reflection, and the pure-Perl `_qnorm()` (Acklam's approximation, also `~1e-9`) that `cohen_d` used for its interval. - Measured against `mpmath` at `mp.dps = 60`, bisecting `erfc(-z/sqrt(2))/2 = p` rather than calling a library inverse, worst relative error over conf.level 0.8 to 0.9999: - | | worst relative error | | |---|---|---| | 0.302 and earlier, at these nine sites | 1.8e-9 | 8.0e6 ulp | | R 4.6.1 `qnorm` | 4.8e-16 | 2.1 ulp | | SciPy 1.18.0 `norm.ppf` | 1.5e-16 | 0.7 ulp | | 0.303 | 8.2e-17 | 0.4 ulp | - Seven orders of magnitude, and the result is closer to the truth than either reference. Nothing in the existing suite failed: every one of these functions is cross-validated against R at tolerances that had room for `1e-9` in them, because they had to. What changed is that they no longer need it — `t/distributions.R.scipy.t`'s `glm` agreement check went from `5e-9`, with its comment explaining that the gap was the design, to bit-identical. - The bug this exposed: `0.95` is a `double`: - With the critical value down to 0.4 ulp, the next-largest error in a confidence limit turned out to be the confidence level. Eleven functions declared their default as NV conf_level = 0.95; - and `0.95` in C is a literal of type `double`, so on a long-double or `__float128` perl what lands in the `NV` is `0.95` rounded to 53 bits — 2.2e-17 from the `NV` nearest `0.95`. `p = 1 - (1 - conf_level)/2` inherits all of it, the normal density at the 95% point is `0.0584`, and the critical value comes out 3.8e-16 wrong: about 1700 ulp of a long double, and the single largest error in a 95% interval on those builds. `fisher_test` and `binom_test` had dodged this for releases by parsing the literal through perl's own number parser; that idiom is now the `NV_CONF_95` macro and all thirteen sites use it. - This was invisible while the critical value itself was only good to `1e-9`. It is the reason `t/qnorm.crit.R.scipy.t` freezes `p` as well as `conf.level`: `1 - (1 - 0.999)/2` has to round, and it rounds differently at each NV width, so a shared expected value is only meaningful if every build is asked the same question. - `t/qnorm.crit.R.scipy.t` is 166 tests — the value itself against R 4.6.1, SciPy 1.18.0 and the 60-digit arbiter, and then the same value recovered out of seven functions' reported intervals at six confidence levels each. It also pins what `inverse_normal_cdf()` alone returns and asserts it is nowhere near the tolerance, so the file cannot pass by accident if the routing is ever undone. [Result keys are dotted everywhere] - Every key of every returned hash now uses R's dotted spelling. Previously the module answered to both conventions and it was impossible to guess which: `p_value` from nine functions but `p.value` from two, `conf_int` from four but `conf.int` from three, `prop_test` returning `conf.int` and `p_value` in the same hash, and `shapiro_test` returning both spellings of its own p-value. - | was | is | |---|---| | `p_value` | `p.value` | | `conf_int`, `conf_level` | `conf.int`, `conf.level` | | `null_value`, `null_value_name`, `statistic_name` | `null.value`, `null.value.name`, `statistic.name` | | `estimate_x`, `estimate_y` | `estimate.x`, `estimate.y` | | `group_stats`, `std_err` | `group.stats`, `std.err` | | `odds_ratio`, `risk_ratio`, `risk_diff` (+ `_ci`) | `odds.ratio`, `risk.ratio`, `risk.diff` (+ `.ci`) | | `risk_exposed`, `risk_unexposed` | `risk.exposed`, `risk.unexposed` | | `n_pos`, `n_neg`, `auc_se`, `auc_ci` | `n.pos`, `n.neg`, `auc.se`, `auc.ci` | | `n_active`, `n_inactive`, `n_top`, `active_count`, `enrichment_factor`, `rie_min`, `rie_max` | the same with dots | | `n_risk`, `n_event`, `n_censor` | `n.risk`, `n.event`, `n.censor` | | `exp_coef`, `lr_stat`, `lr_p_value`, `lr_df`, `loglik_null` | `exp.coef`, `lr.stat`, `lr.p.value`, `lr.df`, `loglik.null` | | `has_na`, `old_coords` | `has.na`, `old.coords` | | `p_adjust` (`dunn_test`'s column) | `p.adjust` | - **This breaks code that reads the old names, and it is the only change here that does.** `$r->{p_value}` is now `$r->{'p.value'}` — note the quotes, which a dotted key requires. `table_one`'s output column is renamed with everything else. - Argument names are untouched. They are a different surface, where the underscore spelling exists on purpose as a synonym for the dotted one, so `conf_level => 0.99` and `'conf.level' => 0.99` both still work everywhere they did, and `wilcox_test(..., conf_int => 1)` still asks for an interval. The duplicate stores are gone: `shapiro_test` and `kruskal_test` each wrote their p-value twice under the two spellings and now write it once. [Decorative comment rules removed] - 1,495 lines that were nothing but a comment marker and a run of dashes, and 380 comments whose text trailed off into one, are gone from `LikeR.xs`, `lib/Stats/LikeR.pm` and `t/`. `/* ===== WILCOXON EXACT NULL DISTRIBUTIONS =====` is now `/* WILCOXON EXACT NULL DISTRIBUTIONS`, which is what the house style in `CLAUDE.md` asks for anyway. Nothing after `__DATA__` was touched, no POD region, and no line holding a `/*` or `*/` was deleted, so no comment can have been left unterminated. The suite is unchanged at 28,485 tests either side of it. 0.302 2026-08-22 CDT - `sum`, `min`, `max`, `mean`, `sd`, `var`, `quantile`, `cor` and `cov` are between 2.5 and 10.7 times faster, and four of the nine are now faster than R. `merge` is between 1.9 and 5.3 times faster and now beats R's `merge` at every size measured, and `uniq` is about four times faster at every size measured. Nothing any of them returns changes, except where it was wrong: five bugs turned up while the measurements were being taken. One had been hanging the interpreter, one had been segfaulting, and three had been quietly returning the wrong frame for any input built on a tied array. - The distribution had also been building without an optimizer, which is where a third of this came from and none of the credit does. - At `n = 1,000,000`, one column of `rnorm` values, best of seven, measured by `scale.pl` against `scale.R` and `scale.py` on the same machine: - | | 0.301 | 0.302 | R 4.6.1 | NumPy 2.5.2 | |---|---|---|---|---| | `sum` | 4.07 ms | 1.37 ms | 0.795 ms | 0.202 ms | | `min` | 4.29 ms | 1.51 ms | 1.06 ms | 0.160 ms | | `max` | 4.25 ms | 1.50 ms | 1.06 ms | 0.160 ms | | `mean` | 4.07 ms | 1.37 ms | 1.59 ms | 0.204 ms | | `sd` | 8.25 ms | 2.39 ms | 2.80 ms | 0.960 ms | | `var` | 8.13 ms | 2.40 ms | 2.79 ms | 0.962 ms | | `quantile` | 160 ms | 14.9 ms | 20.6 ms | 15.0 ms | | `cor` | 20.5 ms | 8.24 ms | 6.43 ms | 4.67 ms | | `cov` | 27.8 ms | 8.75 ms | 4.82 ms | 4.79 ms | - The ratios are better below a million, where the column still fits in cache. At `n = 100,000`, `sum` and `mean` are 5.1 times faster than 0.301, `sd` and `var` 5.2, `quantile` 8.0, and `sum` overtakes R. - What is left between this and NumPy is not a constant factor waiting to be found. A perl array is an array of `SV*`; NumPy holds a contiguous `double[]` and runs SIMD over it. 1.2 ns per element is about four cycles — load the pointer, load the flags, branch, load the `NVX`, add — and that is near the floor for walking a million SVs. Closing it would need a packed-buffer type holding `NV*` directly, which is a different module rather than a tuning pass. [The module was being built without an optimizer] - `compile.sh` ran `perl Makefile.PL OPTIMIZE='-Wall'`. `OPTIMIZE` *replaces* perl's own `$Config{optimize}` rather than adding to it, so that was not "the usual flags plus a warning switch" — it was the whole optimize line, and every build it made, including the one `make install` then installed, was compiled at `-O0`. `test.all.perls.pl` defaulted to the same string, so the entire support matrix was built and timed that way too. - Both say `-O2 -Wall` now. On the same `LikeR.c`, summing `1e6` NVs takes 4.10 ms at `-O0` against 2.63 ms at `-O2`, and `cor` 25.5 ms against 12.9 ms. Read the rest of this section as what was left after the compiler was allowed to do its job. - `t/build.optimize.t` reads both scripts and fails if either hands `Makefile.PL` an `OPTIMIZE` with no `-O` in it. There is no way to ask a loaded `.so` what it was compiled at, so the check has to be on the scripts; neither ships, so the file skips outside the repository. [Reading a column: one `av_fetch` per element] - All nine functions read their input by calling `av_fetch` once per element. That call is out of line and repeats, for every element, a bounds check, a negative-index fixup and a magic dispatch that a straight walk of `AvARRAY` already knows the answers to. On a million-element column it was about half of what `sum` cost and more than half of `cor`. - Walking the block directly is only sound while nothing can move it, and `SvNV` can move it: on an SV carrying get magic it runs `mg_get`, on a reference it may run an overloaded `0+`, and on a string that is not a number it raises a warning that a `$SIG{__WARN__}` handler catches. Each of those is perl code, and perl code may push to the very array being walked, which reallocs the block and leaves the walk holding freed memory. - So the fast path is decided per element, not per array. `sv_plain_nv` takes an SV only when it already holds a number and carries neither magic nor a reference, which makes `SvNV` a register read that calls nothing; everything else — a hole, a string, a reference, a tied element — goes back to `av_fetch`, which is both correct and where the block pointer gets re-derived. A tied array is refused whole, before the scan starts. - The scans live in their own functions and are marked `LIKER_NOINLINE`, which is not decoration. Inlined into the caller, a scan shares a stack frame with the `av_fetch` loop that finishes whatever it would not take; the accumulators then have to stay addressable across that call, so the scan spills them to memory every iteration instead of keeping them in registers. Summing `1e6` NVs: 2.77 ns/element for the `av_fetch` loop being replaced, 2.25 with the scan inlined into it, 1.29 with the scan compiled on its own. There is no portable spelling of `noinline`, so it is annotated where the compiler has one and left empty where it does not, which gives back the inlined speed and never a wrong answer. - `min` and `max` get a scan each rather than sharing one that tracks both ends. The running extreme is a loop-carried dependency, each compare-and-move waiting on the previous element's result, so carrying the end the caller will not return costs a second chain for nothing: 1.79 ms against 1.35 ms on `1e6` NVs. [`var`, `sd` and `cov`: Welford to two passes] - Welford's recurrence updates the running mean with `mean += delta / count`. That divide sits in a loop-carried dependency chain, so each element waits out the previous one's latency — 5.34 ns/element against 2.70 for two passes. - Two passes is also what R computes. `src/library/stats/src/cov.c` takes a first mean, refines it with a second pass (`tmp = tmp + sum(x - tmp)/n`, the `MEAN` macro) and only then sums the squared deviations about the refined mean. Expanding that third pass about the *unrefined* mean leaves `sum((x-m)^2) - (sum(x-m))^2/n`, so carrying the sum of the deviations alongside the sum of their squares gives R's answer in one pass fewer. That is the Chan-Golub-LeVeque correction, so it is no less accurate than Welford on an ill-centred column. - `mean` is deliberately not changed: it remains the single-pass `sum/n`, which is not what R's `mean` computes and never was. - Reading the arguments twice is visible only to a tied array, whose `FETCH` now runs once per pass. [`quantile` orders only the statistics it reads] - `quantile` sorted the whole column and then indexed into it. Type 7 reads at most two order statistics per probability — `x[j]` at `j = floor((n-1)p)`, and `x[j+1]` where the interpolation weight is non-zero — which for the five default probs is ten values out of a million. - `nv_select_multi` places just those. It takes the middle index first and recurses into the two ranges that index separates, costing `O(n log k)` rather than the `O(n k)` a left-to-right sweep of selects would: each level of the recursion touches each element at most once, and the index list halves at every level. Past `k = n/2` distinct indices a full sort is the cheaper way to get the same array, and that is what still happens there. - This is also what R does — `quantile.default` hands its indices to `sort(partial=)`, and NumPy reaches `np.partition` for the same reason. On `1e6` random doubles the ten order statistics of the five default probs cost 19.2 ms this way against 69.1 ms for the full introsort. - The probabilities are parsed before the column is read now, rather than after it is sorted, so an invalid `probs` vector no longer costs a pass over the data first. [`cor` and `cov`: sorts that inline their comparison] - Spearman ranks its input before correlating it, and `rank_data` did that through `qsort` with an out-of-line comparator; Kendall's tau-b sorted three times the same way. That comparator call cannot be inlined, and for a sort of a few thousand elements it is most of what the sort costs — the module's own `nv_sort` takes 210us where `qsort` takes 355us on 5,000 NVs, which is why it exists. - `LIKER_DEFINE_SORT(T, PFX, LESS)` generates that introsort for a given struct and ordering: median of three left at the midpoint so both partition scans have a sentinel, recursion into the shorter side and a loop on the longer so the stack stays `O(log n)` deep, a heapsort fallback past `2*floor(log2 n)` levels so a median-of-three killer cannot reach `O(n^2)`, and one insertion sort over the nearly-ordered remainder. It is `kw_sort`, the Kruskal-Wallis sort, made generic. `rank_data` and Kendall's `(x, y)` sort are generated from it; Kendall's third sort, a plain `NV` column, calls `nv_sort` directly. - glibc's `qsort` is also a mergesort that allocates a scratch buffer the size of the array it is given. Peak RSS for a Spearman correlation of two `2e6` columns falls from 455 MB to 424 MB — exactly the 32 MB it was allocating to sort `2e6` sixteen-byte records. - At `n = 200,000`: `cor(..., 'spearman')` 66.9 ms to 27.2 ms, `cov(..., 'spearman')` 65.9 ms to 26.7 ms, `cor(..., 'kendall')` 70.5 ms to 44.7 ms. Small repeated calls gain the same factor: ten thousand `cor(..., 'spearman')` calls over twelve elements, the shape `agg` and `group_by` produce, 6.4 ms to 3.1 ms. [`cor` read an array of arrays once per column] - `cor(\@matrix)` extracted one column at a time, and for each column it walked every row — `av_fetch` on the outer array, `SvRV` to reach the row, then `av_fetch` for the cell. The outer array was therefore walked `ncols` times over, and every one of the `ncols * nrows` cells paid a fetch on it that the column before had already paid. - The pre-validation pass that already checks every row is an array reference now keeps the row `AV*`s it resolves instead of discarding them, and the extraction transposes in a single pass over the rows through `av_extract_or_nan` — the same reader the vector branch uses, so a short row, a hole, an `undef` and a non-numeric cell all behave here exactly as they do there. Outer fetches fall from `ncols * nrows` to `nrows`, and each row's cells are read in the order they are stored rather than one per pass. - At 20,000 rows: `cor` of a 40-column matrix 40.6 ms to 26.7 ms, and its cross-correlation against an 8-column matrix 28.7 ms to 12.4 ms. The transpose keeps one write stream open per column, which is what a wide matrix pays for this, and 2,000 x 300 still comes out ahead — 144.6 ms to 134.8 ms. - A row that was itself a tied array used to read as all-`NA`, because `av_fetch` on one hands back a deferred `PVLV` whose value only arrives when `mg_get` runs: `cor(\@rows)` croaked "standard deviation is 0 in x column 0" on data that was perfectly well defined. `av_extract_or_nan` runs the get magic, so those rows now give the same answer as the untied equivalent. A cell that is a tied scalar was already read correctly and still is. - The zero-variance check that used to be folded into the extraction is now `nv_all_equal`, one function shared with the vector branch, which returns at the first pair of values that differ instead of sweeping the column. On the vector path that replaces two full passes over `x` and `y`, which is where the `cor` row of the table above moves from 10.5 ms to 8.24 ms — the two passes were a fifth of what `cor(\@x, \@y)` cost at `n = 1,000,000`. - `cov` does not gain from any of this: its Pearson path neither sorts nor scans for zero variance, and 8.75 ms is the same figure it read before. What changed there is housekeeping — `Newx` in place of a bare `safemalloc`, and the method string resolved once rather than by three more `strcmp` calls. [`min` over more than 65,535 arguments never returned] - `min` counted its arguments in an `unsigned short int`. `items` is the whole flattened argument list, so `min(@x)` on an array of more than 65,535 scalars wrapped the counter and the loop never reached its bound: my @a = (1 .. 70_000); min(@a); # 0.301: hangs, forever - It is a `Stack_off_t` now, which is what `max` and `sum` already used. `power_t_test` carried the same counter; it takes named pairs, so nothing could reach 65,535 of them, and it is widened anyway. - A regression here is a hang rather than a wrong answer, so `t/hot_path.t` runs the 70,000-argument calls under `alarm` where the platform has one. Without that guard a smoker would sit there instead of failing. [A tied array read as all-`undef`] - `sum(\@tied)` croaked `undefined value at array ref index 0`, and so did every other function here. `av_fetch` on a tied array does not return the element: it returns a mortal `PVLV` that only acquires the element's value once `mg_get` has run on it. The `SvOK` test came first, saw an empty `PVLV`, and reported every element of every tied array as undefined. An ordinary array holding a tied scalar failed the same way. - The slow path runs `SvGETMAGIC` before it looks now. This predates 0.302 and is not fallout from the scan — but the scan is what made it worth finding, because `sv_plain_nv` refuses every magical SV and so sends all of them down that one path. - That slow path then had a bug of its own, and it was worse. `av_slow_at` held the `SV **` that `av_fetch` returned — a pointer *into* `AvARRAY` — ran the get magic through it, and read it again afterwards. `mg_get` runs perl, and the element in `t/hot_path.t` pushes 200 values onto the array being walked, which reallocates that block: both reads after the magic were reads of freed memory. It usually returned the right answer because the freed block usually still held the old pointer, so the test that exists for precisely this case passed on four of the five perls and, on the fifth, only failed when the suite was run four perls at a time and something else had claimed the block — `sum` returned 1085 instead of 10, once. Under `valgrind` it is an "invalid read of size 8" on every perl, every time. - It loads the element into a local before running the magic now. The SV itself does not move; only the array of pointers to it does, so a copy of the pointer stays good where a second read through the slot does not. The same shape — take `SV **` from `av_fetch`, run magic, dereference again — is still present in the `median` and `moment` readers and in three places in `transpose`, and is the same one-line fix in each; those are untouched here. - `uniq` was missed by this pass and fixed later in the same release; see "`uniq`: the same keys, without a perl hash" below. [`cor` returned `NaN` on an off-centre column] - `pearson_corr` computed the correlation from raw cross-products, `(n*sxy - sx*sy) / sqrt((n*sx2 - sx^2) * (n*sy2 - sy^2))`. Each of those three terms subtracts two nearly equal large numbers, so it loses a digit of the answer for every digit by which the mean exceeds the spread. On a column with mean `1e9` and a spread of `0.05` it does not merely lose precision: `n*sx2 - sx*sx` comes out negative, the square root is `NaN`, and `cor` returns `NaN` for an ordinary pair of vectors. At mean `1e6` it was wrong in the fifth decimal place — `0.74945606763640826` against R's `0.74943997817595964`. - It centres first now, as R's `cov.c` does, with the same compensated two-pass correction `var` uses. R reports `0.063454463733801467` for the `1e9` case, and so does this. - `cov`'s Pearson branch was Welford's covariance recurrence, two divides deep in the dependency chain; it is `nv_cov2` now, the two-pass form, shared with the Spearman branch. [A leak on `cor`'s matrix error path] - `cor(\@x, \@y)` with a `y` matrix whose rows are empty croaked "y matrix has zero columns" without freeing the `x` columns it had already extracted. On a 40-column matrix that is 40 buffers and their table per call: 20,000 such croaks grew RSS by 44 MB. The branch's allocations are now released through one macro that every croak path calls, and the row tables and the scratch row are freed as soon as the last value has been read out of them rather than at the end. - The `nrows < 2` guard was missing the same frees, but it is unreachable — a one-row matrix has a constant column, so the zero-variance croak above always gets there first. It frees anyway. - Seven `croak` formats in `cor` passed a `size_t` to `%lu` with no cast, which on Windows is a 64-bit argument read through a 32-bit conversion. They use `%" UVuf "` with a `(UV)` cast now, as the rest of the file does, and `cov`'s one — cast to `unsigned long`, so well defined but still narrowing on Win64 — went with them. Every message is unchanged wherever `unsigned long` is 64 bits, which is everywhere the suite runs. [Unchanged: `cor` and `cov` still do not see an overloaded object as a number] - `cor` and `cov` decide what is numeric with `looks_like_number`, which is false for any reference, an overloaded `0+` included, so a column of overloaded objects reads as all-`NA`: `cov` returns `NaN` and `cor` reaches its zero-variance croak. `sum`, `min` and `max` do honour the overload, through `SvNV`. The two families genuinely disagree. That predates 0.302 and is a different question from how a column is walked, so `t/hot_path.t` pins the current behaviour rather than changing it quietly. [`merge`: a hash join that stopped building a perl hash] - At `n = 300,000` a `merge($df, $df, how => 'inner', on => 'id')` over a six-column HoA took 575 ms. It takes 109 ms. Best of ten, one process per measurement, pinned to one CPU, against `scale.R` and `scale.py` on the same machine: - | `n` | 0.301 | 0.302 | R 4.6.1 | pandas 2.x | |---|---|---|---|---| | 1,000 | 0.471 ms | 0.245 ms | 0.46 ms | 0.68 ms | | 10,000 | 5.45 ms | 2.60 ms | 3.35 ms | 0.88 ms | | 100,000 | 138 ms | 33.3 ms | 80.0 ms | 2.95 ms | | 300,000 | 575 ms | 109 ms | 155 ms | 9.83 ms | - Under callgrind the 30,000-row join fell from 358M instructions to 170M. Three things were paying for that. - The right-hand index was a perl `HV` keyed by the join key: an `HE`, a shared `HEK` and an `IV` SV per distinct key, so three allocations per right row whenever the key is unique, which is exactly what a join on an id column is. It is the same open-addressed table over a key arena that `drop_duplicates` already interns into (`dd_ctx`) — no SV, no HEK, one `memcpy` of each distinct key — so the two functions now share their definition of "the same key" as well as their code for deciding it. `mg_key` writes straight into that arena instead of making four `sv_catpvn` calls per key column into a scratch SV, and a one-column join writes the cell's bytes bare: with a single field there is nothing for a length prefix to disambiguate, and on an integer id the prefix and its two separators were about as many bytes again to hash and to compare. - It emitted each output row as it found it, so the row count was not known until the join had finished and every output column grew by `av_push`, being reallocated its way up to the answer. The probe now records `(left row, right row)` pairs into a flat list and builds nothing; only when that list is complete — so the count is exact — is each column allocated once at its final size and filled straight into `AvARRAY`, the way `filter` already builds its columns. The list costs two `SSize_t` per output *row* against roughly one SV per output *cell*, so it is a small fraction of the result on any join wide enough to care, a cross join included. `mg_emit` is gone; `mg_column` and `mg_build` replace it, and the cross join goes through the same two stages as every other `how`. - `mg_cell` called `av_fetch` once per cell, which was 7% of the call: a join reads every key cell of both frames and every cell of every column it keeps. It reads the block directly now, through the same `av_at` the other frame functions use. - What is left is the result itself. A 300,000-row join of two six-column frames is 3.3 million output SVs, and at 109 ms that is 33 ns each — allocate, copy a cell into it, store it. pandas is not doing the same work: its columns are contiguous typed buffers and a join is a `take` over them. [`uniq`: the same keys, without a perl hash] - `uniq` on a million-element column of `rnorm` values took 0.80 s, against R's `unique()` at 0.019 and pandas' `pd.unique()` at 0.029. It takes 0.16 s. Same method as the tables above — one process per measurement, pinned to one CPU, the fastest of seven, `plot.scaling.pl` against `scale.R` and `scale.py` on the same machine — except that the `0.301` column is the 0.301 *code* built at `-O2`, not the released build, so this is the rewrite on its own and not the optimizer again: - | `n` | 0.301 code | 0.302 | R 4.6.1 | pandas 2.2.3 | |---|---|---|---|---| | 1,000 | 0.294 ms | 0.069 ms | 0.009 ms | 0.020 ms | | 10,000 | 3.19 ms | 0.825 ms | 0.092 ms | 0.184 ms | | 100,000 | 41.2 ms | 9.71 ms | 1.16 ms | 2.30 ms | | 300,000 | 174 ms | 39.2 ms | 3.71 ms | 5.59 ms | | 1,000,000 | 801 ms | 160 ms | 19.4 ms | 28.6 ms | - The slope was never the problem. Fitted over the ladder by `scaling.slopes.tsv`, all three are linear — 1.12 for `uniq`, 1.11 for R, 1.04 for pandas, and 1.14 for `uniq` before the rewrite — so this was a constant factor from the start, and none of it was the hashing. - `SvPV` on an `NV` is a `%.15g` that has to be redone every time. Walking a million `NV` SVs costs 2 ms; walking them and calling `SvPV` on each costs 268 ms. That is ten times R's whole runtime for the same call, and it is the floor the old code could not get under however few distinct values there were: a column of a million elements holding ten distinct values still took 0.226 s. Worse, `SvPV` leaves the rendered buffer on the caller's own SV without ever reusing it (`sv_2pv_flags` re-renders an `NV` on the next pass regardless), so asking a large numeric column for its distinct values grew that column by tens of megabytes, for good. `nk_num_pv` — already written for `drop_duplicates`, about four times faster and only taken where the answer is provably the same — renders into the XSUB's own stack buffer and touches nothing. - The `seen` hash was a perl `HV`. An `HE` and a copied `HEK` per distinct key is ~72 MB of scattered small allocations at a million of them, and the table rehashes at every doubling on the way there. Going from ten distinct values to a million added 0.47 s of pure insert cost. It is `dd_ctx` now, the same open-addressed slot array over a key arena that `drop_duplicates` and `merge` intern into, presized from the element count. - It hashed each key twice, once for `hv_exists` and again for `hv_store`. Collapsing that into one `hv_fetch(..., 1)` is the obvious repair and is the wrong one: the lvalue fetch mints an SV per key, and on the million distinct doubles it measured 0.919 s against the pair's 0.756 — while also reading `AvARRAY` directly, so the comparison flatters it. That is why the pair had survived. One hash of one key into one open-addressed probe is what the arena gives instead. - What the rewrite does *not* do is compare doubles by value. That is how R and pandas get the factor of six to eight they still have, and it is a different answer: `0.1 + 0.2` and `0.3` are two doubles that print the same, so they are one value to `uniq` and two to `unique()`. `uniq` is documented to compare the way `eq` and `List::Util::uniq` do, so it renders every element and compares the text; that rendering pass is essentially all of the distance that is left. Scalar context builds no result list at all and takes about a fifth off again. - `uniq(\@tied)` croaked `undefined value at array ref index 0`. It was reading elements through `av_fetch` and testing `SvOK` on the `PVLV` that comes back, which is the bug "A tied array read as all-`undef`" describes above. `uniq` was simply not one of the functions that pass reached. It reads through `av_slow_at` now, like the rest. [`filter`: measured, and left alone] - `filter` is not faster, and the reason is worth writing down so it is not looked for again. At `n = 300,000` on a five-column HoA, `col('x') > 0` keeping half the rows costs 22.5 ms, of which the predicate is 1.3 ms. The other 21.2 ms is allocating and filling the 750,000 SVs of the result, which the documented contract requires: an HoA input, or any `hoa` output, builds fresh arrays and fresh cell values. That is 19 ns per numeric cell and 44 ns per string cell, and the difference between the two is one `malloc` for the string's buffer. - Copy-on-write is the obvious way out of that `malloc` and is not available. `newSVsv` is `newSVsv_flags(sv, SV_GMAGIC|SV_NOSTEAL)`, which does not pass `SV_COW_SHARED_HASH_KEYS`, so its string copies are never copy-on-write; and `SV_DO_COW_SVSETSV`, the flag pair a plain `my $b = $a` gets, is defined as `0` unless `PERL_CORE` is set — perl's own `sv.h` says "the core is safe for this COW optimisation, XS code on CPAN may not be". Reaching past that by hand was measured anyway and is slower in both directions: `filter` 23.0 ms against 28.3 ms, `merge` 103 ms against 126 ms. For a string as short as a data frame's usually are, the `malloc` costs less than what COW puts in its place — an out-of-line `sv_setsv`, the `CowREFCNT` increment, and the read-only flip on the source. The finding is recorded at `flt_cell_copy` rather than only here. - The one lever that does move it is the copy itself. Sharing the surviving cells instead — one refcount bump, which is what `drop_duplicates` does for exactly this reason — was prototyped and measured at 13.6 ms against R's 14.4 ms. It is not done, because `filter` documents the opposite and someone may be relying on it; changing that is a decision about the interface, not a tuning pass. - Two things did change. `flt_num` takes a bare `IOK`/`NOK` cell without calling `SvGETMAGIC` or `looks_like_number` — neither can say anything about an SV that already holds a number, and `looks_like_number` is out of line — which is most of what the predicate pass costs on a numeric column. And the tied-frame bugs below. [A tied frame read as all-`undef`, and one that segfaulted] - `filter`, `merge` and `drop_duplicates` all walk a column or a row through `AvARRAY`, which is why they are fast and which is wrong for a tied array: its elements do not exist until `FETCH` has run, and `AvARRAY` on one is not the block they live in — on an array that has never held a real element it is a null pointer. `av_fetch` is the way in, and what it returns for a tied element is a mortal `PVLV` that only acquires the value once `mg_get` has run on it, so an `SvOK` or `SvROK` test placed before that says "undef" or "not a reference" for every cell and every row of the frame. - This is the same pair of mistakes `sum(\@tied)` made, above, and it predates 0.302 in all three functions. Between them they were wrong in four distinct ways: tie my @x, 'TiedArray'; @x = (1, -2, 3); tie my @y, 'TiedArray'; @y = (10, 20, 30); filter({ x => \@x, y => \@y }, col('x') > 0); # 0.301: { x => [], y => [] } -- an empty frame, whatever the predicate # 0.302: { x => [1, 3], y => [10, 30] } merge({ id => \@x }, { id => [1, 3], w => ['a','b'] }, how => 'inner', on => 'id'); # 0.301: { id => [], w => [] } -- no key matched, because every key read undef # 0.302: { id => [1, 3], w => ['a', 'b'] } drop_duplicates({ k => \@x, v => \@y }); # 0.301: { k => [undef], v => [undef] } -- every row identical, so one survived # 0.302: { k => [1, -2, 3], v => [10, 20, 30] } drop_duplicates([ map { { k => $_ } } 1 .. 3 ]); # with the AoH itself tied # 0.301: segmentation fault - The crash was `_aoh_key_union`, which indexed `AvARRAY` with no bounds check and no test for a tied array at all; on a tied AoH that is a null dereference on the first row. An outer join was the quietest of the four — with no key matching, it returned the two frames as disjoint halves, well-formed and entirely wrong. - There is one reader now. `av_at` returns an element as it is — block read where that is sound, `av_fetch` where the array is tied, and the pointer derived per element rather than hoisted, because copying a cell can run perl and perl can push to the array being walked. `av_ref_at` adds the `mg_get` and the "is it a reference to the right thing" test for a row; `av_row_keep` is how a row gets into a result, handing back the caller's own reference for a plain array and a fresh reference to the same row for a tied one, since the `PVLV` is mortal and stays bound to the tie. `dd_cell` runs the get magic before it looks at the cell. None of it is measurable: `drop_duplicates` is within 2% of 0.301 on all three shapes at 30,000 and 300,000 rows, and `merge`'s numbers above are with it. - One consequence is visible and is documented under `drop_duplicates`: a tied HoA column's cells cannot be *shared* into the result, because the tie has no cell SV to share, so they are copied. It is the only place the "what survives is shared" rule cannot hold, and copying is the only reading of it that can be true. - A tied *frame* hash — the outer hash of a HoA or HoH, rather than a column or a row inside it — is still not supported by any of the three. All of them read its shape and its columns through `HeVAL(hv_iternext(...))`, which for a tied hash is not the value. That is a larger change than this one and is not made here; `t/tied.frames.t` says so where it stops. [Tests] - Three new test files, and the suite is 129 files and 26,331 tests. - `t/var_sd_cov.R.t` (162 tests) cross-validates `sum`, `min`, `max`, `mean`, `var`, `sd`, `cov` and `cor` against R 4.6.1 over ten columns chosen to separate the algorithms rather than to be representative: an arithmetic ladder, a constant column, `n = 2`, mixed signs, and the column with mean `1e9` and spread `0.05` that breaks a naive variance and did break the old `cor`. The expected values are frozen literals with their provenance in the file header; the generator, `t/var_sd_cov.R.R`, is committed beside it, and the test itself never calls R. - Every value in the corpus is a dyadic rational, so the vectors are the same numbers at every NV width, and the corpus is rebuilt in perl rather than pasted in — the `n`, `sum`, `min` and `max` columns are what pin the perl construction to R's. Those frozen numbers are printed `%.40g`, the double's exact decimal expansion, because 17 significant digits round-trip a double back to a double and no further: a long-double perl reads `100000000004.83398` as a different number from the `100000000004.833984375` R had, which was enough to fail an exact comparison on two of the five perls. The tolerance on the exact columns is two ulps of whatever NV the running perl was built with, found by bisection at run time rather than hardcoded, because perl's string-to-NV conversion is not correctly rounded on a long-double build — `perl-5.12.5` reads that literal one ulp short. - Every case also runs through a tied array, which cannot take the scan and has to reach the same answer through `av_fetch`. That is the assertion that the fast path and the slow path are the same function. - `t/hot_path.t` (135 tests) covers how the input is read rather than what comes out: the 70,000-argument counter; every kind of element the scan must refuse — IV, UV above `IV_MAX`, NV, string, `0+` overload, tied element, and one magical element in the middle of an otherwise plain column; holes and `undef` in all three of their documented behaviours; and an element whose get magic pushes 200 values onto the array it lives in while that array is being walked. `cor`, `cov` and `var` are checked for invariance under a location shift of `1e3`, `1e6` and `1e9`, which is the property the old correlation failed, and `quantile`'s partial sort is checked against a full sort at every rank on a column with ties. - `t/tied.frames.t` (81 tests) covers `filter`, `merge` and `drop_duplicates` over frames built on tied arrays and tied row hashes: every output shape, every `how`, every `keep`, single and composite keys, an empty tied frame of each shape, and the sharing rule on both sides of its one exception. Every case runs the same call over tied input and over an identical plain copy and requires the two answers to be equal — a fixed expected value would pin the plain answer as well, which `t/filter.t`, `t/merge.t` and `t/drop_duplicates.t` already do, and what has to be pinned here is that the two routes agree. `t/merge.t`'s existing reference join — plain Perl, run over all six input/output shape combinations — is what checks that the rewritten join still means what it meant. - `./test.all.perls.pl` passes on all five local perls — `5.10.1`, `5.12.5` (long double), `5.42.3`, `5.44.0` and `5.44.0-quadmath` — with no warnings on any of them, and `LikeR.c` compiles clean under `-std=c99 -O2 -Wall` against every one of their `CORE` directories. 0.301 2026-08-21 CDT - there are numerous additions of `restrict` keywords, which may or may not improve speed [kruskal_test] - `kruskal_test` cross-validated against R 4.6.1's `stats::kruskal.test()` and SciPy 1.18.0's `TestKruskal`, driven by those suites' own cases rather than by cases invented here. Six bugs are fixed: five in how the arguments and the data are read before any ranking happens, and one in the chi-squared tail, which reaches every function that uses it. The test now needs a third of the memory and runs in under half the time, and it returns the same answer twice in a row, which it did not before. - Everything below is checked in the new `t/kruskal_test.R.scipy.t` (870 tests), whose expected values are frozen literals with their provenance recorded in the file header; it needs no R and no Python to run. The generator that produced them, `t/kruskal_test.R.scipy.R`, is committed beside it. There had been no `t/kruskal_test.t` at all — the only coverage was six assertions in `t/01.t` on the single Hollander & Wolfe example, and none of the six bugs would have shown up in it. The full suite is 125 files and 25,951 tests, and `./test.all.perls.pl` passes on all five local perls — `5.10.1`, `5.12.5` (long double), `5.42.3`, `5.44.0` and `5.44.0-quadmath` — with no warnings on any of them. - NaN was ranked instead of dropped: - `looks_like_number` is true for `NaN`, so a `NaN` went into the ranking. R treats `NaN` as `NA` and `complete.cases(x, g)` removes it before `rank()` ever sees it: - | data | was | R 4.6.1 | |---|---|---| | `c(1,2,3,4,5,6)` with one `NaN`, n = 7 | `H = 4.5` | `H = 3.8571428571428577` | | `1:24` with one `NaN` | `H = 17.28` | `H = 16.5` | - It also handed `cmp_nv3` a comparison that is never true for any pair involving the `NaN`, which leaves `qsort` without the strict weak ordering it is entitled to — the same defect `wilcox_test` had fixed in 0.298, where the comment describing it is still in the file. `+Inf` and `-Inf` are neither `NA` nor `NaN` to R and a rank test has no trouble with them, so they are still kept and ranked; that is now pinned rather than incidental. - The bad value propagated: `table_one`'s `_t1_cont_p` hands its groups straight to `kruskal_test`, so a single `NaN` among nine observations was reported at `p = 0.0273` where dropping it gives `0.0439`. - A group with no data inflated the degrees of freedom: - The hash-of-arrays form counted every key in `k`, including a key whose array was empty or whose every element had been dropped. On `{a => [1,1,1], b => [2,2,2], c => []}` that gave `df = 2` and `p = 0.0820849986238988` for what is a two-group problem with `df = 1, p = 0.025347318677468304`. Such a group was already skipped when forming the statistic and when building `group.stats`; only `df` still counted it. - R refuses this case outright — `all groups must contain data` — and so does `kruskal_test` now, because the alternative is to test the groups that do have data under a `df` that counts one that does not. R's order of checks came with it: it filters each group, refuses an empty one, and only then counts what is left, so `{a => [], b => []}` is `all groups must contain data` and not `not enough observations`. SciPy takes the other side of this and returns `NaN` with a `SmallSampleWarning`; the divergence is recorded in the test file rather than papered over. The `x`/`g` form cannot reach any of this — it mints a group id the first time an observation survives the filter — so nothing changes there. - Group labels were truncated at a NUL and lost their UTF-8 flag: - The `x`/`g` path read the label with `SvPV_nolen` and then took `strlen` of it. Perl strings are counted, not NUL-terminated, so `"a\0X"` and `"a\0Y"` collapsed into one group: `kruskal_test([1..6], ["a\0X","a\0X","a\0Y","a\0Y","b","b"])` came back as two groups with `df = 1, H = 2.4` instead of three with `df = 2, H = 4.571428571428573`. Both paths also copied the label's bytes while dropping perl's UTF-8 flag when storing it into `group.stats`, so a label outside latin-1 came back as mojibake and the two input paths disagreed with each other about labels inside it. The length now travels with the string and carries the flag in its sign, which is `hv_store`'s own convention, so a label comes back `eq` to what went in. Dropping the `strlen` also drops a pass over every label. - A trailing named argument read past the argument stack: - The named-argument loop took `ST(arg_idx + 1)` without checking that there was one, so an odd argument list read one slot past the top of the stack — and what it found there changed which branch ran: `kruskal_test(\%h, 'x')` came back complaining that `'h'` cannot be mixed with `'x'`/`'g'`, because `x_sv` had been assigned whatever was past the end. `binom_test`, `chisq_test`, `fisher_test`, `wilcox_test`, `var_test` and `prcomp` all guard this; `kruskal_test` was the one that did not. It now croaks `odd number of named arguments`. - An infinite chi-squared statistic gave no p-value: - `get_p_value` short-circuits a statistic at or below zero and otherwise goes to `igamc`. `+Inf` is neither, so it reached the continued fraction, where the first `1/d` is `1/Inf = 0` and then `del = 0 * Inf` is `NaN` — an overwhelmingly significant result reported as no result at all. R's `pchisq(Inf, df, lower.tail = FALSE)` is `0`, and so is this now. `NaN` in gives `NaN` out, as R does, rather than running the continued fraction to its full 10,000-iteration safety bound first. - This is reachable from `kruskal_test`. When a sample has no variation at all the tie correction is `(n^3 - n)/(n^3 - n)`, and once `n^3` is past `2^53` the subtraction of `n` is lost from one side or the other, so an inexact zero is divided by an exact zero. R has the same problem and returns `+Inf`, `-Inf` or `NaN` depending on which way `n` rounded: `NaN` at `n = 250000`, `-Inf` at `300000`, `NaN` at `400000`, `+Inf` at `500000`, `NaN` at `750000`, `+Inf` at `1000000`, `NaN` at `1500000` and `+Inf` at `2000000`. `kruskal_test` now agrees with it on all eight. Below that the correction is exact and both give `NaN`, which the corpus pins. `get_p_value` is shared, so `chisq_test`, `prop_test`, `mcnemar_test`, `friedman_test`, `cmh_test`, `logrank_test` and `coxph` get the same fix. - Three times less memory, twice the speed: - The ranking no longer goes through `RankInfo` and `rank_and_count_ties`. `kruskal_test` wants per-group rank sums, not the ranks themselves, so it sorts a 16-byte `(value, group)` pair and adds each tie block's averaged rank straight into the group sums, instead of storing an `NV` rank per observation for a second pass to read. - It also no longer calls `qsort`. glibc's `qsort` is a mergesort that allocates a scratch buffer the size of the whole array — measured on glibc 2.39 as a `VmHWM` of 116 MB going to 230 MB across one sort of a 114 MB array — which was half of the function's peak memory, and its comparison goes through a function pointer that cannot be inlined. In its place is a median-of-three introsort that recurses on the smaller partition and loops on the larger, so the stack stays `O(log n)`, with a heapsort fallback past a depth of `2*floor(log2(n))` so an adversarial input cannot drive it to `O(n^2)`, and insertion sort for short runs. Sorting the same five-million-element array takes 0.375s against `qsort`'s 1.01s and allocates nothing. The third change is the group-label array on the `x`/`g` path, which was sized at one pointer per *observation* to hold one per *group* — 40 MB at `n = 5e6` to hold three pointers — and now grows on demand. - At `n = 5,000,000` over three groups, measured as `VmHWM` either side of the call: - | | 0.3 | 0.301 | |---|---|---| | peak memory | 228 MB (47.8 B/obs) | 76 MB (15.9 B/obs) | | `kruskal_test(\@x, \@g)` | 1.20 s | 0.557 s | | `kruskal_test(\%h)` | 1.11 s | 0.467 s | - Sorted, reversed, all-equal, organ-pipe and median-of-three-killer inputs all stay under 0.32s at `n = 2e6`, which is what the depth limit is there for. The sort is checked against an independent pure-Perl implementation of the whole test over 748 structured cases — those shapes at every n either side of the insertion-sort threshold — and 49,712 random ones. - The same input now gives the same answer: - `H` moved by up to `1.2e-14` between runs on identical data. Nothing was random: the sum of `R_i^2 / n_i` walked the groups by group id, and on the hash-of-arrays path an id is minted in `hv_iternext` order, which is perl's per-process hash order. Equal values were also left in whatever relative order the sort happened to leave them, which came from the same place. - The sort now orders by value and then by group, which makes it a total order, and the `k` terms of the sum are ordered before they are added — smallest first, which is the better-conditioned direction as well as a canonical one. `k` is the number of groups, not the number of observations, so it costs nothing next to the ranking. `H` is now bit-identical to R on all 37 corpus cases in all four call forms, and stays so across 60 runs under `PERL_PERTURB_KEYS=1`. [Documentation] - `kruskal_test` gains two sections: what happens to non-numeric, undefined, `NaN` and infinite elements and to a group left with no data, and what the returned fields are — `statistic`, `parameter`, `method` and the p-value under both `p.value` and `p.value` from R's `htest`, plus the `size` and `mean` sub-hashes of `group.stats`, which are computed over the same observations the statistic used. 0.3 2026-08-16 CDT [shapiro_test] - `shapiro_test` rebuilt against R 4.6.1's `src/library/stats/src/swilk.c` — AS R94, Royston (1995) — driven by R's and SciPy's own test suites rather than by cases invented here. Four bugs are fixed, one of them a case R's regression suite tests for by name, and the statistic is now more accurate than R's own on a sample whose values dwarf its spread. - Everything below is checked in the new `t/shapiro_test.R.scipy.t` (146 tests), whose expected values are frozen literals with their provenance recorded in the file header; it needs no R and no Python to run. The generators that produced them, `t/shapiro_test.R.scipy.R` and `t/shapiro_test.R.scipy.py`, are committed beside it. The full suite is 124 files and 25,081 tests, and `./test.all.perls.pl` passes on all five local perls — `5.10.1`, `5.12.5` (long double), `5.42.3`, `5.44.0` and `5.44.0-quadmath` — with no warnings on any of them. - The p-value could come back negative: - R's `tests/reg-tests-1b.R` contains exactly one `shapiro.test` assertion, and it is this: stopifnot(shapiro.test(c(0,0,1))$p.value >= 0) - `shapiro_test([0,0,1])` returned `-4.6648135328131477e-15`. At `n = 3` the p-value is `6/pi * (asin(sqrt(W)) - asin(sqrt(3/4)))` and `W` has an exact floor of `3/4` that `c(0,0,1)` sits on, so the subtraction lands on zero from whichever side the constants round to; R clamps the result at 0 and this module did not. It is the same defect SciPy fixed as gh-18322. The clamp is in, and `asin(sqrt(3/4)) = pi/3` is now carried to NV width rather than R's 15 digits, so the same case comes out at `+4.2e-16` before clamping instead of below zero. - W and the p-value were only good to nine digits: - The expected normal order statistics that AS R94 weights the sample with came from `inverse_normal_cdf()`, which is Moro's approximation and good to about `1e-9`. They go straight into `W`, so nine digits there is nine digits in the answer — where R reports sixteen. Against the values SciPy pins in `TestShapiro`, every one of them annotated upstream as *"reference values generated using R shapiro.test"*: - | SciPy case | W was | W is | R 4.6.1 | |---|---|---|---| | `test_basic` x1 | `0.900472879324135` | `0.900472879317561` | `0.90047287931756` | | `test_basic` x2 | `0.959026945965277` | `0.959026946032345` | `0.95902694603234` | | `test_basic2` x4 | `0.834666275331324` | `0.834666275318169` | `0.83466627531817` | - and the p-values with them: - | SciPy case | p was | p is | R 4.6.1 | |---|---|---|---| | `test_basic` x1 | `0.0420895752342124` | `0.0420895752222577` | `0.04208957522226` | | `test_basic` x2 | `0.524597929157127` | `0.524597930470668` | `0.5245979304707` | | `test_basic2` x4 | `0.000913490482316994` | `0.000913490481812984` | `0.000913490481813` | - Moro's value is still the starting point, but Newton against the `erfc`-based normal CDF finishes it. The loop stops as soon as another pass could not move the answer, so a `double` build pays for one refinement and only the wider NVs pay for a second. Across a 180-sample sweep over normal, uniform, exponential, log-normal, Cauchy, tied, tiny-scale and grid data at every n from 3 to 5000, the worst remaining disagreement with R is `2.0e-15` in `W` and `2.6e-12` in the p-value — the latter is not sloppier arithmetic but the same last bit amplified, since the p-value is a function of `log(1 - W)` over a sigma of about `0.6`. - 1 - W was formed by subtracting from 1: - The p-value depends on `log(1 - W)`, and `W` runs to within `1e-5` of 1 on a large normal sample, so computing `W = b^2/ssq` and then `1 - W` throws away exactly the digits the p-value is made of. R does not do this — its `swilk.c` forms `w1 = (ssassx - sax) * (ssassx + sax) / (ssa * ssx)` directly and says so in a comment — and now neither does this module. - More accurate than R when the values dwarf their own spread: - R's `swilk.c` divides the sample by its range but never centres it, so `1e9 + noise` loses most of its significant digits before `W` is ever formed. SciPy filed the same complaint from the other end as gh-14462 and works around it by subtracting the median; this module now does that too, which costs one subtraction per value. - Measured against a 60-digit `mpmath` evaluation of AS R94 on the identical doubles, over `1e6 + noise` and `1e9 + noise` at every n from 3 to 5000: - | | worst relative error in W | worst in the p-value | |---|---|---| | R 4.6.1 | `1.9e-8` | `2.6e-7` | | `shapiro_test` | `1.0e-15` | `1.4e-13` | - On well-conditioned samples the two still agree to the last few ulp, so this is a divergence only where R has already lost the digits. `t/shapiro_test.R.scipy.t` asserts the invariance rather than R's number there, and records why at that section. - Faster as well: - The sort now goes through the module's own introsort rather than `qsort()`, whose comparator the compiler cannot inline; the order statistics are generated for half the sample and mirrored, since the weights are antisymmetric; and ten `pow()` calls became Horner evaluations. `pow()` is a `__float128` call on a quadmath perl. - | n | was | is | |---|---|---| | 10 | 0.83 µs | 0.78 µs | | 100 | 4.78 µs | 5.05 µs | | 1000 | 79.3 µs | 49.8 µs | | 5000 | 578 µs | 459 µs | - `n = 100` is the one size that got slower: at that length the accurate quantiles are most of the work and there is not enough sorting to pay for them. That trade was taken deliberately. - One documented value was wrong: - The hash printed under `shapiro_test` in this README, in `read.me.pod` and in the module's own POD showed `statistic 0.960870680168535` and `p.value 0.589650577093106` — the pre-fix numbers, and for `[1..19]` while the example above them calls `shapiro_test([1..5])`. It now shows what `[1..5]` actually returns, `0.986762155447719` and `0.96717393596804`, which is R's `shapiro.test(1:5)` to the last digit R prints. [quantile] - Interpolation ran between order statistics that were equal: - R's type 7 interpolates only when the index falls strictly between two order statistics *that differ* — `index > lo & x[hi] != qs` in `quantile.default`. This module always evaluated `(1 - g) * x[j] + g * x[j+1]`, which does not return `v` when both sides are `v`. On a two-valued sample at `n = 999` it reported `0.99999999999994` for `1`, and on the 602 identical values R's PR#16672 was filed about it failed to return that value at every prob — which is the monotonicity failure the PR is about. Both now match R exactly. - Probabilities a hair outside [0, 1] were refused: - A probability arrived at by arithmetic rather than written down can land just outside the interval. R allows `100 * .Machine$double.eps` of overshoot and clamps to the endpoint — its PR#17891, `quantile(0:1, 1+1e-14) == 1` — where this module raised an error. It now clamps within the same allowance and still errors on anything further out. R's constant is used rather than `NV_EPSILON` on purpose: it is part of what the function *accepts*, so a long-double or `__float128` build must not reject a `probs` vector R takes. - Faster: - The sort was `qsort()` with a function-pointer comparator; it is now the same introsort `shapiro_test` uses. Ordering 5000 NVs costs about 61,000 comparisons, and paying for an indirect call on every one of them is most of what a sort of that size costs. - | n | was | is | |---|---|---| | 100 | 3.2 µs | 2.7 µs | | 1000 | 52.5 µs | 22.8 µs | | 10,000 | 1.07 ms | 0.72 ms | | 100,000 | 13.7 ms | 8.8 ms | - Both fixes and the sort are covered by the new `t/quantile.R.t` (197 tests, or 205 under `EXTENDED_TESTING`), built on the two assertions R's own suite makes about `quantile` — that `quantile(x, ((1:n)-1)/(n-1))` recovers `sort(x)`, and that it equals the type-7 interpolation computed by hand off the sorted sample — run over seven input shapes chosen to break a quicksort (sorted, reversed, organ pipe, two-valued, tie ladder, sawtooth) at every n either side of the insertion-sort threshold and the recursion depth limit, plus PR#16672 and PR#17891 verbatim and 79 frozen R value tables. Its generator, `t/quantile.R.R`, is committed beside it. [Documentation] - Illustrations for three more functions, drawn by `t.test.plots.pl` and `skew.kurtosis.plots.pl`, both committed: - **`t_test`** gains six: what the estimate, the standard error and the null distribution are and which area of it the p-value is; how `conf.int` is the estimate plus or minus a t quantile and how `conf.level` sets that quantile; the three `alternative`s side by side with the region each counts and the interval that goes with it; `p.value` as a function of `mu`, crossing `1 - conf.level` exactly at the two bounds of `conf.int`; paired, `var_equal` and Welch on the same data, with the Welch degrees of freedom as the two spreads separate; and two distributions separating with the interval retreating from `mu` as the p-value falls. - **`skew`** gains a left-tailed, a symmetric and a right-tailed sample against the same `N(0, 1)` curve, with the mean and median of each, which is what the sign of the statistic is reporting. - **`kurtosis`** gains a flat-shouldered, a normal and a heavy-tailed sample with the tails behind each drawn out, since it is the tails and not the peak that the statistic is measuring. 0.298 2026-08-12 CDT [wilcox_test] - A rewrite of `wilcox_test` against R 4.6.1, driven by R's and SciPy's own test suites rather than by cases invented here. It brings the function up to the exact conditional inference R gained in 4.6.0, fixes six bugs — two of which returned confidently wrong p-values on the *default* code path — and adds the Hodges-Lehmann estimate and confidence interval, `digits.rank`, and the Edgeworth series. - Everything below is checked in the new `t/wilcox_test.R.scipy.t` (3,242 tests), whose expected values are frozen literals with their provenance recorded in the file header; it needs no R and no Python to run. The full suite is 120 files and 23,149 tests, and `./test.all.perls.pl` passes on all five local perls — `5.10.1`, `5.12.5` (long double), `5.42.3`, `5.44.0` and `5.44.0-quadmath` — with no warnings on any of them. - Exact p-values are now computed when there are ties: - R 4.6.0 added exact (conditional) inference in the presence of ties, via Torsten Hothorn's implementation of the Streitberg-Röhmel shift algorithm; R's `doc/NEWS.Rd` announces it and `tests/reg-tests-1d.R` records the consequence at its degenerate one-sample cases: *"For R >= 4.6.0 warnings for exact with ties are gone."* Before that, ties ruled out an exact p-value and both R and this module fell back to the normal approximation with a warning. - `wilcox_test` now does what R does. When ties are present the null distribution is the conditional one given the observed ranks, and the same holds for zero differences in the signed-rank test. The warnings are gone with them. - This changes published answers on tied data, including R's own documented examples: - | case | was | is (R 4.6.1) | |---|---|---| | `?wilcox.test` man-page data, `wilcox_test(\@x, \@y)` | `0.13291945818531886` | `0.12990538872891813` | | the `airquality` Ozone example (`W = 127.5`) | `1.2080783e-04` | `6.1087351888e-05` | | `wilcox_test([1,2,2,3], [4,5,5,6], exact => 1)` | `0.02842953599879653` + a warning | `0.028571428571428571` | | `wilcox_test([1,1])` | `0.34577858615116` | `0.5` | | `wilcox_test([4,3,2], [3,2,1], paired => 1)` | `0.14891467317876567` | `0.25` | - Two further consequences are worth knowing about. V itself changes when zero differences are present, because the exact test ranks `|x - mu|` over every observation and only afterwards drops the ranks belonging to the zeroes, where the approximation drops the zeroes first and ranks what is left: `wilcox_test([-1, 0, 1])` gives `V = 2.5` exactly and `V = 1.5` with `exact => 0`. R's two branches differ in exactly the same way. And degenerate inputs that used to be fatal now return a result, as they must for `tests/reg-tests-1d.R` line 332 to pass: `wilcox_test([0])` gives `V = 0`, `p = 1`, and so does `wilcox_test([0,0,0,0,0])`, which SciPy pins as `test_all_zeros_exact`. - If you need the old numbers, `exact => 0` still asks for the approximation and is unchanged. - The exact upper tail was returning zero, on the default path: - `p_greater` was computed as `1 - CDF(q - 1)`. That subtraction cancels away every significant digit once the true p falls below `NV_EPSILON`, and then returns a flat `0`. It did not take a contrived input to reach: two perfectly separated samples of 30 apiece are inside the automatic exact branch, no `exact => 1` required. - | m = n | was | is | R 4.6.1 | |---|---|---|---| | 20 | `7.2544192875e-12` | `7.2544445519e-12` | `7.2544445519e-12` | | 25 | `7.8825834748e-15` | `7.9107286024e-15` | `7.9107286024e-15` | | 30 | `0` | `8.4556169461e-18` | `8.4556169461e-18` | | 49 | `0` | `3.9250145965e-29` | `3.9250145965e-29` | - Both tails are now summed directly. That alone is not enough for the rank-sum table, whose Gaussian-binomial recurrence is built with subtractions, so far up the support a count of `1` is the difference of numbers around `C(m+n, n)` and has already been rounded into noise. The table is folded about its centre before summing, so only well-conditioned entries are ever touched — the same thing R's `pwilcox()` does when it folds `q` about `m*n/2` and flips `lower_tail`. - The signed-rank tail was accurate to `n = 49` by luck (`1 - 2^-49` is exactly representable) and reached `0` from about `n = 53`; forcing `exact => 1` on `n = 120` returned `0` where R gives `1.5046327690525337e-36`, and now returns it too. - `int m * n` overflowed, and said the samples were identical: - `exact_pwilcox` took `int m, int n` and computed `int max_u = m * n`. For two separated samples of 50,000 that wraps negative, every statistic looks out of range, and the function returns `1.0`: wilcox_test([1 .. 50000], [50001 .. 100000], exact => 1); # p = 1 - Signed overflow is also undefined behaviour, so a different optimiser was entitled to do something else entirely. Sizes and indices in the exact distributions are `size_t` now, the multiplications are checked for wrap before they happen, and a table that would need more than 16 million cells is refused outright with a message naming `exact => 0` rather than attempted. - NaN was ranked instead of dropped: - `NaN` is `NA` to R, and R drops it. `looks_like_number` accepts it, `d == 0.0` is false for it, so it went into the rank buffer — and `cmp_nv3` returns `0` for every comparison involving it, which leaves `qsort` without the strict weak ordering the C standard entitles it to. - The visible symptom is R's own regression case, `tests/reg-tests-1d.R` line 3546, which asserts that a paired test is unaffected by pairs whose difference is `Inf - Inf`: - | | was | is (and R) | |---|---|---| | `1:5` vs `4*(0:4)` | `V = 1`, `p = 0.125` | `V = 1`, `p = 0.125` | | the same with `+Inf` appended to both | `V = 1`, `p = 0.0625` | `V = 1`, `p = 0.125` | | the same with `-Inf` and `+Inf` on both | `V = 2`, `p = 0.046875` | `V = 1`, `p = 0.125` | - `NaN` — in either sample, and however it arises — is now dropped with the other missing values. `±Inf` is not missing and is kept, since a rank test has no trouble with it; SciPy's `test_gh_11355b` pins five cases of that and they all agree. - An empty `y` ran a different test: - `wilcox_test([1,2,3], [])` fell through to the one-sample branch and returned a signed-rank result, silently answering a question nobody asked. It croaks now, with R's message. - `mu` was likewise unvalidated: `mu => Inf` or `mu => NaN` turned every difference into a non-number and produced a confident answer from the wreckage. Both croak now, as they do in R. - A dying `$SIG{__WARN__}` handler leaked the rank buffer: - The warnings in `wilcox_test` were emitted while the `RankInfo` and difference buffers were held as raw pointers. A `__WARN__` handler that dies — or `warnings FATAL` at the call site — longjmps straight past the `Safefree`. Under valgrind, 500 iterations of the ties path with such a handler lost 95,616 bytes in 498 blocks. Every allocation now goes through `Newx` plus `SAVEFREEPV`, the idiom `chisq_test` in the same file already used, so it is released by the save stack however the call unwinds. The same 500 iterations now report `definitely lost: 0 bytes`, as does a sweep over every croak path and every branch of the function. - New: `conf.int`, and a Hodges-Lehmann estimate: - R has returned a distribution-free confidence interval and a point estimate since PR#1150 in 2001, and `tests/reg-tests-1a.R` has guarded them ever since with Hollander & Wolfe's published numbers. `wilcox_test` now computes both, by all four of R's routes — the exact interval from the order statistics of the Walsh averages or the pairwise differences, the exact interval conditional on the observed ranks when there are ties, and the asymptotic interval from a root search: my $r = wilcox_test(\@y, \@x, paired => 1, conf_int => 1); # $r->{estimate} == -0.46 # $r->{conf.int} == [-0.786, -0.010] # $r->{conf.level} == 0.9609375 - Those are Hollander & Wolfe (1999) 2nd ed., pp. 40 and 53, to the digit. So are the two-sample values from pp. 111 and 126: estimate `-0.305`, interval `(-0.76, 0.15)`. - The level a rank test can actually deliver is a step function of the data, so `conf.level` reports what was achieved rather than echoing what was asked for — `0.9609375` above, not `0.95`. `conf.level`, `tol.root` and R's alpha-doubling search for a level the data can support (with its *requested conf.level not achievable* warning) all behave as R's do. - New: `digits.rank`, `edgeworth`, and more of R's result fields: - `digits.rank` rounds each value to a given number of significant digits before ranking, so that ties are decided on the rounded values. R's man page recommends it because tie detection is an exact `==` on floating point, and its own worked example shows `(4:2)/10` against `(3:1)/10` — three differences that ought to be `0.1` and are three different doubles. Ported from R's `fprec()`, half-to-even rounding included. - `edgeworth => 1, 2, 3` adds up to three Edgeworth correction terms to the normal approximation, the refinement R 4.6.0 reaches through its integer `correct`. It is ignored on the exact path, and — as in R — ignored when there are ties, or when the signed-rank test dropped a zero, because the series is derived for untied ranks. - The result hash gains `statistic.name` (`"W"` or `"V"`, as R prints), plus `null.value` and `null.value.name`, and `estimate` / `conf.int` / `conf.level` when an interval was asked for. - Three deliberate differences from R: - Each is asserted in the test file, so that changing one later is a choice rather than a drift. - 1. `correct` is a boolean here. R 4.6.0 turned its `correct` into an integer `0:3`, in which numeric `0` still applies the continuity correction and only `FALSE` removes it — so in R, `correct = 0` and `correct = FALSE` are different tests. Keeping that would mean `correct => 0` no longer meaning "off", which is what it means for every other flag in this module. `correct` stays a boolean, and R's `correct = k` is `correct => 1, edgeworth => k`. 2. A zero variance is reported, not propagated. With `exact => 0` and every observation tied there is nothing to divide by. R divides anyway and returns `NaN`; this warns and returns `p = 1`. The default path no longer reaches it at all, since the exact test handles all-tied data. 3. An all-tied interval does not raise. R's one-sample code warns and hands back a `NaN` interval at level `0`; its two-sample code warns and then dies inside `uniroot` with *missing value where TRUE/FALSE needed*. We give the one-sample answer in both places. - There is one place where this module is simply more accurate than R. R's exact p-values on tied data come from a density it normalises entry by entry; `wilcox_test` sums the integer permutation counts and divides once. For the worst case in the corpus — an 11-against-12 tied rank sum whose p-value is exactly `4/676039` — this returns the correctly rounded double and R is `1.2e-11` high. Checked against exact rational arithmetic, and recorded in the test file rather than papered over. - Testing: - `t/wilcox_test.R.scipy.t` takes its cases from the references' own suites: - R's `tests/reg-tests-1a.R` (the PR#1150 Hollander & Wolfe intervals), `reg-tests-1b.R` (the Wolfgang Huber `wilcox.test(1, 2:60)` case, and the check that the asymptotic estimate does not move with `alternative`), `reg-tests-1d.R` (the six degenerate one-sample calls and the `±Inf` identities), and the man-page examples whose printed output is pinned in `tests/Examples/stats-Ex.Rout.save`. - SciPy 1.17.1's `TestMannWhitneyU`, whose header reads *"All magic numbers are from R wilcox.test"* — `cases_basic`, `cases_continuity`, `cases_9184`, `cases_2118`, `test_tie_correct`, `test_exact_U_equals_mean`, `test_gh_11355b` and the 30-against-20 asymptotic cases — and `TestWilcoxon`'s `test_accuracy_wilcoxon`, `test_wilcoxon_tie`, `test_onesided`, `test_exact_pval`, `test_exact_p_1`, `test_all_zeros_exact` and `test_symmetry_gh19872_gh20752`. - A 663-case sweep generated by `t/wilcox_test.R.scipy.R`, committed next to the test, crossing four data shapes against every alternative, `exact` state, `correct` state, `mu` and `conf.int` setting. - Beyond the file, 960 further randomised calls were compared against R 4.6.1 and agree everywhere except the three divergences above. - One lesson from getting that to pass on every NV width is worth recording: the corpus data has to be exactly representable. Whether two values tie decides which branch runs, and `1.6 - 2 - 0.5` does not land on the same value in a `double`, an x87 `long double` and a `__float128`. A corpus of one-decimal values passed on the default perl and failed on `perl-5.12.5` and quadmath with a *different statistic*, not merely a different last digit. Every generated value is now a whole number of quarters or of 1024ths. For the same reason the asymptotic interval, which is only ever pinned down to `tol.root`, is generated at `tol.root = 1e-12` rather than freezing wherever Brent's method happened to stop on one machine. - A compiler-warning audit of `LikeR.xs` for `-Wint-conversion`, `-Wimplicit-int`, `-Wreturn-mismatch` and `-Wdeclaration-missing-parameter-type`, and a pass tightening integer types that can only hold a count or a flag. No behaviour changed: the full suite (116 files, 18,546 tests) passes, and every function touched was diffed call-for-call against a build of the previous release, with `dnorm`, `pnorm`, one- and two-sample `ks_test`, `fisher_test`, `auc` and the set operations re-checked against R 4.6.1 and found bit-identical. - All four of those warnings were already clean, and stay clean on a `double`, a `long double` and a `__float128` build. Two of them cannot be tested with the GCC most systems still default to: `-Wreturn-mismatch` and `-Wdeclaration-missing-parameter-type` are GCC 14 additions — where they are errors rather than warnings — and GCC 13 rejects both as unrecognized options, so a check that appears to pass on 13 has really only skipped them. [Dead code removed] - Turning the audit up to `-Wextra` found two branches that could never run, both of them a test for negativity on a value whose type is unsigned: - 1. `r_pow_di` takes `unsigned int n`, so its `if (n < 0) return 1.0 / r_pow_di(x, -n);` was unreachable — a leftover of R's `R_pow_di`, which takes a signed exponent. All three callers (in `K2x`, for the exact one-sample Kolmogorov-Smirnov distribution) pass a non-negative exponent, so the unsigned parameter is the correct one and the reciprocal branch simply goes. 2. `hoa2aoh` casts `HvUSEDKEYS` to `U32` and then clamps with `if (ncols < 0) ncols = 0;`. [Types narrowed to what they can actually hold] - Eighteen `int`s that only ever hold 0 or 1 became `bool`, a convention the file already followed in some 219 other places; each was confirmed by reading every call site rather than by name. The flag parameters of `ft_pnhyper`, `K2l`, `c_dnorm`, `c_pnorm`, `c_pnorm_both`, `set_multiplicity` and `roc_split`, the `is_cat` field of `AnFac`, the `lower_pos` and `frac_low` locals of `auc`, `auroc`, `roc` and `bedroc`, and the return types of `mg_key` and `psmirnov_exact_test`. Several of these were already being handed a `bool` by their callers — `dnorm`'s and `pnorm`'s `log` and `lower` options, for instance — so only the helper signatures were behind. `c_pnorm_both`'s loop counter became `unsigned int`. - Two that look like flags and are not: `c_pnorm_both`'s `i_tail` is three-valued, and `set_multiplicity`'s `gimme` carries a Perl `G_*` context value. Both stay `int`. - Also six coefficient tables in `c_pnorm_both` written `const static double`, which puts the storage class after the qualifier and draws `-Wold-style-declaration`; they are now `static const double`. [runif argument validation, and every warning names its function] - `runif` accepts its arguments either positionally or by name, and decided which was which by asking whether the current argument was a string *and* whether another argument followed it. A key at the end of the list therefore failed the second half of that test and fell through to the positional branch, where it was read as a number: `runif(5, 'min')` took `SvNV("min")`, which is 0, silently set `min = 0`, and returned five values. The only sign anything was wrong was perl's own `Argument "min" isn't numeric`, which does not say which function provoked it. `runif(5, bogus => 1)` went the same way, taking `bogus` as `min` and `1` as `max`. Every sibling that parses named arguments — `rbinom`, `binom_test`, `fisher_test`, `dnorm`, `pnorm` — rejects both of those. - `runif` now does too. A string argument is treated as a key when it is not a number, which is decidable from the key alone, so a dangling or misspelled key is an error instead of a silent coercion; a numeric string is still positional, so `runif("9")` is unchanged. Named values are checked for numerichood before use, which is what keeps perl's unattributed warning from being the diagnostic. - `n` is also range-checked now. It was read straight through `SvUV()`, so `runif(-1)` wrapped to 2**64-1, `av_extend()` read that back as a negative `SSize_t`, and perl died with `panic: av_extend_guts() negative count (-2)` -- which names neither the function nor the argument at fault. A negative or over-large `n` now croaks and says so. Non-integer `n` still truncates toward zero, as R's `runif()` does, and `runif(0)` still returns an empty list. - Separately, three warnings did not name the function emitting them, unlike every other warning in the file: one in `ks_test` (the 1-sided exact 1-sample case falling back to asymptotic) and two in `wilcox_test`'s signed-rank branch (exact p-value abandoned for ties, and for zeroes). All three now carry the prefix their siblings already had. The one warning left deliberately bare is the `warn("%s", m)` in the uninitialized-value catcher, which re-emits somebody else's warning verbatim and must not add to it. [Argument-stack indices are now Stack_off_t] - `-Wextra` reported 58 `-Wsign-compare` warnings, and 33 of them were one idiom: an index declared `size_t`, `unsigned`, `unsigned int` or `unsigned short int` and then compared against `items`. `items` is neither of those — XSUB.h's `dITEMS` declares it `Stack_off_t items = (Stack_off_t)(SP - MARK)`, a *signed* type, because it is a stack-pointer difference. Every one of those comparisons was converting the signed side to unsigned. - The indices are now `Stack_off_t` themselves, which is the type they are compared against: 25 declarations across 23 functions — `binom_test`, `ks_test`, `wilcox_test`, `write_table`, `max`, `runif`, `quantile`, `mean`, `mode`, `sum`, `sd`, `uniq`, `var`, `t_test`, `median`, `matrix`, `fisher_test`, `power_t_test`, `var_test`, `dnorm`, `value_counts`, `prcomp` and `pnorm`. That is a retype, not a cast: writing `(size_t)items` at each comparison would silence the warning just as well, but it would be wrong the day `Stack_off_t` widens, which is exactly what it exists to allow. `t_test`'s index was `unsigned short int`, which drew no warning at all — integer promotion made the comparison signed — and was the same latent mistake regardless. - `Stack_off_t` arrived in perl 5.39.2 and this distribution supports 5.010, so the preamble now carries a shim typedef guarded on `PERL_STACK_OFFSET_DEFINED`, the macro perl.h defines next to the typedef. On 5.10.1 and 5.12.5 neither the macro nor the type exists and the shim supplies `I32`, which is what the stack offset was on every perl before that. - The 33 warnings are gone, 25 remain, and no warning category increased — verified by compiling the before and after trees and diffing the warning sets. The remaining 25 are unrelated signedness pairs (`size_t` against `ssize_t`, `IV` against `size_t`, `STRLEN` against `ssize_t`) and are left alone. The full suite passes on perl 5.10.1 and 5.12.5, the two builds that depend on the shim, as well as on 5.42.3, 5.44.0 and 5.44.0-quadmath; and 94 calls covering all 23 retyped functions — positional and named forms, bare lists against arrayrefs, `write_table`'s emitted bytes, and the odd-argument and unknown-argument croaks that this index arithmetic drives — produce identical output before and after. [NV was being computed at double precision on wide builds] - Every libm call in `LikeR.xs` was written bare — `sqrt(x)`, `log(x)`, `lgamma(x)` — and C has no type-generic ``. Those functions take a `double`, so on a perl built with `-Duselongdouble` or `-Dusequadmath` every one of them converted the `NV` down to 53 bits of mantissa, computed there, and converted the result back. Nothing warned and nothing failed to compile; the answers were simply less accurate than the perl running them. On perl-5.12.5 (`long double`), `sd(1..5)` returned exactly the double-rounded `sqrt(2.5)`, 9.5e-17 away from the value perl's own `sqrt` gives. - All 412 of those calls now go through `nv_*` macros that paste on the suffix for the width `NV` actually is: none for `double`, `l` for `long double`, `q` for `__float128`. The 80 `isnan`/`isinf`/`isfinite` calls became `nv_isnan`, `nv_isinf` and `nv_isfinite`, which classify by comparing against `NV_MAX` — the largest finite `NV` — rather than calling libm at all. The C99 macros could not be kept: where a platform does not provide the type-generic versions, `isfinite()` is a plain `double` function, and narrowing a large-but-finite long double into it reports the value as infinite rather than merely rounding it. Perl's own `Perl_isnan`/`Perl_isinf`/`Perl_isfinite` were used up to 0.298 and could not be kept either: on every perl before 5.22 those route through a `Perl_fp_class()` block in `perl.h` that has never compiled — the macro is written with an empty parameter list and compares against `FP_CLASS_*` names no `` defines. That block is dead code wherever Configure finds `isinf()`, so it is invisible on Linux and glibc, and live on illumos/Solaris, where it broke the 0.298 build outright. - The long-double row is conditional. The `l` variants are C99 but some libms — the thinner BSD ones especially — do not ship the whole set, so `Makefile.PL` link-tests all twenty as a unit and defines `LIKER_HAVE_LONG_DOUBLE_MATH` only if every one resolves; otherwise the build falls back to the `double` functions, which is exactly what it did before and so cannot regress. `__float128` needs no probe: `` and `-lquadmath` come with the quadmath perl itself, and the built object was checked with `nm` — it references `lgammaq`, `expq`, `sqrtq` and no double-width libm symbol at all. - Accuracy on the long-double build, measured against values that are exact in binary or known in closed form: `sd(1..5)` is now bit-identical to perl's `sqrt(2.5)`, and `fisher_test([[3,1],[1,3]])` moves from 1.5e-16 to 6.4e-18 relative error against the exact 17/35. The remaining 6.4e-18 is an accuracy floor in that function's own summation, not a width problem — the `__float128` build lands on the same figure. - This costs time where the wide math is software-emulated: the suite takes 352s on the quadmath perl, against 67s when it was quietly running on hardware doubles. The other four perls are unaffected. [The build ran itself twice, and clobbered its own Makefile doing it] - `make` had to be run twice or the `.so` came out stamped with the wrong version and refused to load. The cause: ExtUtils::MakeMaker scans the directory for `*.PL` files to run during the build, and `dev.Makefile.PL` — a local convenience wrapper, not part of the distribution — looks like one. It was being run mid-build as `perl dev.Makefile.PL dev.Makefile`, and since it calls `WriteMakefile()` it overwrote the real `Makefile` with its own: no `DEFINE`, no probed C99 flag, and a different `VERSION`. The second `make` then rebuilt from that. `PL_FILES => {}` turns the scan off; nothing here is generated by a `.PL` file. - The version half was a stale literal: the checked-in `Makefile.PL` pinned `VERSION => "0.28"` while `lib/Stats/LikeR.pm` had moved to 0.298, and `XSLoader::load()` passes `$VERSION` to a `.so` compiled with `-DXS_VERSION` from that literal. It now reads `VERSION_FROM => lib/Stats/LikeR.pm`. One `make` after `perl Makefile.PL` is enough again, and the non-quadmath builds are about a third faster for not doing the work twice. [Portability: Solaris, the BSDs, and vendor compilers] - The C99 flag is now probed instead of guessed. `Makefile.PL` was selecting `-std=gnu99` on any compiler whose name matched `/\b(?:g?cc|clang)\b/`, and `$Config{cc}` is plain `cc` for Oracle Studio on Solaris and for aCC on HP-UX — both of which reject that flag outright, so the build failed there before it compiled a line. Each candidate is now trial-compiled and the first that works wins: `-std=gnu99`/`-std=c99` for gcc and clang, `-xc99=all` for Studio, `-qlanglvl=extc99` for AIX `xlc`, `-AC99` for HP-UX, and nothing at all for a compiler already in C99 mode. MSVC is skipped outright, since it warns rather than errors on switches it does not know and would make the probe settle on a no-op. - Two things that would have failed to compile off Linux are gone. `` and its `strcasecmp` — POSIX-only, absent on MSVC — are replaced by a small `str_ieq_ascii()`, which also drops the locale dependency: `tolower()` under a Turkish locale maps `I` outside ASCII, which should never decide whether `"TRUE"` matches `"true"`. And bare C99 `restrict`, used on 151 pointers here, now has an `#ifdef` mapping it to `__restrict` on MSVC and `__restrict__` on older gcc, and defining it away where no spelling exists, rather than losing the annotation. - `LikeR.xs` also compiles clean under strict `-std=c99` with no GNU extensions, which is the closest available local proxy for a vendor compiler. [Dead code: sample()'s private PRNG] - A splitmix64 generator sat at the top of the file under a comment promising a PRNG stream separate from `Drand01()`, seeded lazily from `/dev/urandom` with a `time()^PID` fallback. None of it was true: no seeding code was ever written, no caller ever existed, and its state started at a fixed 0, so had anything called it the "random" sample would have been the same sequence in every process. `sample()` draws from `Drand01()` and always did, which is the behaviour that is wanted — `srand($seed)` governs it the way `set.seed()` governs R. The generator and its comment are removed. [Tests] - Two files, 273 assertions, and both were checked against a deliberately broken build rather than merely observed to pass. - `t/nv_width.t` fails if the math width ever comes undone. Its sharp assertion needs no tolerance at all: `sd(1..5)` must be the identical NV to perl's `sqrt(2.5)`, which holds on any width and breaks the moment a `double` gets in the way. It is width-adaptive rather than skipped on a `double` perl, computing the NV epsilon of the running build instead of assuming one. - `t/scale.keywords.t` covers `scale()`'s string options — `"mean"`, `"sd"`, `"none"`, `"true"`, `"false"`, `""` and their case variants — which had no coverage at all: `t/01.t` passes only the numeric forms. Expected values come from R 4.6.1 `base::scale()` at `options(digits=17)` and are frozen in the file, so it needs no R at run time. Deleting the case fold from `str_ieq_ascii()` fails 11 of its assertions; usefully, all 11 are the "off" spellings, because an unmatched string falls through to `SvTRUE` and still means "compute it", so `"MEAN"` would keep working while `"NONE"` flipped. That is recorded in the file so the section is not trusted for more than it proves. - The suite is 118 files and 18,819 tests, passing on perl 5.10.1, 5.12.5, 5.42.3 (threaded), 5.44.0 and 5.44.0-quadmath, with no compiler warnings on any of them. 0.297 2026-08-10 CDT - https://www.cpantesters.org/cpan/report/260534ea-9474-11f1-8ca2-bfb68deea6df bug fix 0.296 2026-08-09 CDT - fixed CPAN bug: https://www.cpantesters.org/cpan/report/fcf32c68-75a5-1014-bc87-8fe0d10910fe - write_table.announce.t ran its child perl through -e, which cannot carry double quotes or shell metacharacters on Windows; the child program now goes in a file - chisq_test now matches R 4.6.1 bit-for-bit on the statistic across 170 randomized cross-check cases, and the full suite (116 files, 18,546 tests) passes. - Bugs found and fixed in LikeR.xs - 1. A 1×k or k×1 table returned df = 0, p = 1 — no test at all. R collapses a single-row/column matrix to a vector and runs goodness-of-fit (if (min(dim(x)) == 1L) x <- as.vector(x)); now so does this. [[10,20,30]] went from X²=0, df=0, p=1 to X²=10, df=2, p=0.006738. 2. Yates' label was attached even when the correction was zero. R only says "with Yates' continuity correction" when min(0.5, |O−E|) > 0. A table sitting exactly on its expectation, and every zero-margin table, were mislabelled. 3. Yates was computed per cell instead of as R's single whole-table min(0.5, abs(x-E)) — equal in theory on a 2×2, not always in the last bits. 4. No input validation. Negatives, infinities, NaN, strings and undef were silently coerced to 0 and produced garbage or NaN; all-zero data returned NaN; a single element returned df = 0. All now croak with R's wording. Ragged array rows and 2D hash rows with mismatched column keys were silently zero-filled — now fatal. 5. Uniform expectation used n/k instead of R's n * (1/k), and sums were accumulated in a plain NV where R uses a long double. Together these put the statistic 1–2 ulp off R on most inputs; both fixed (ct_acc_t). 6. Hash input was read in Perl's randomized key order, so which row a malformed hash got blamed on was a coin toss. Rows and columns are now sorted, as fisher_test already does. 7. Segfault on sparse arrays (av_fetch returns NULL for a hole) — this one I introduced during the rewrite and caught before finishing; guarded by ct_av_get - Three of those cross-checks compared the statistic to R's printed value relatively, and on the tables in question R's value is not a statistic. Where a 2×2 has all four |O−E| equal, Yates' min(0.5, |O−E|) cancels every corrected residual, so the exact statistic is 0 and the exact p is 1; what R prints there — 1.4515367733818938e-24 for [[1573,3],[4,0]], 2.9347503914472165e-32 for [[1,2],[3,4]], 7.1842689582627857e-32 for [[1.5,2.5],[3.5,4.5]] — is the leftover of forming E in floating point, the four |O−E| differing in their last bits so that the minimum comes out a hair below the rest. Its size is a property of the NV rather than of the test: a double build reproduces R's digits, and a __float128 build cancels the whole way to 0. Comparing that relatively can only pass on the width R happened to use, and it failed with rel diff = 1 on the quadmath perl and on 5.12.5. Those three cases in t/chisq_test.R.scipy.t now check the statistic against 0 and the p-value against 1 with absolute tolerances of 1e-20 and 1e-11, R's numbers staying in the file as provenance. LikeR.xs is unchanged — the wide-NV answer was the more accurate one. The suite passes on perl 5.10.1, 5.12.5, 5.42.3-thr, 5.44.0 and 5.44.0-quadmath. 0.295 2026-08-08 CDT - bug fix https://www.cpantesters.org/cpan/report/0f13fed6-92f5-11f1-b043-dc326e8775ea - Removed `restrict` where it made no difference, or was potentially dangerous [drop_duplicates, merge, value_counts] - These three decide what counts as the same row, the same join key, or the same value by a cell's Perl stringification, and on numeric columns that one conversion was most of the work they did. - `sv_2pv_flags()` renders an NV with `snprintf("%.*g", NV_DIG, x)`, about 140 ns a cell, and — unlike the IV case, where `SvPOK_or_cached_IV` lets the `SvPV` macros hand back the string perl cached on the SV — it never reuses that PV, so every pass over a column of doubles paid the conversion again. It does leave the buffer behind, which is why keying a frame used to grow the caller's own numeric columns by about 64 bytes a cell, permanently: reading a frame ought to be a read. - `nk_num_pv()` now renders bare integers and bare doubles into the caller's own scratch buffer instead, and leaves the SV untouched. Its double path is `%.15g` about four times faster than the C library's, and taken only where the answer is provably the same: the magnitude is scaled into `[1e14, 1e15)` in `long double` — 64 mantissa bits against the double's 53 — which bounds the scaled value's error under 2e-4, so a fractional part further than 2e-3 from one half rounds exactly as the true value would. About one cell in 300 lands nearer than that and goes back through `SvPV`, as do zero, the non-finite values, `use locale`, an x87 control word left at double precision, and any build whose NV is not an IEEE double. It agreed with the C library's own `%.15g` over 90 million random bit patterns; `t/drop_duplicates.t`, `t/merge.t` and `t/value_counts.t` now group tens of thousands of doubles both ways and require the two answers to match. - Two further changes in `drop_duplicates` alone: - Its interning table started at 64 slots and doubled, so a pass over 10,000 distinct rows rehashed nine times, each one a scattered walk over a table too big for L2. The row count is known before the pass starts and bounds the group count, so it is now used as the hint — capped, so a large frame of few distinct rows does not pay for a slot per row. - An HoA result copied every surviving cell, while AoA and AoH already shared the whole surviving row. It now shares the cells too. **This is a behaviour change.** The frame, and an HoA's column arrays, are still new, so they can be reshaped without touching the input; but assigning *through* a survivor — `$out->{col}[0] = ...` — now writes to the input's cell, exactly as `$out->[0]{col} = ...` always did for AoA and AoH. Clone the result if you need full independence. - Measured on the 10,000-row frame `benchmark.pl` uses (five columns: two doubles, one integer, two strings), on one machine, with only these paths toggled. Time is the median of 25 calls in one process; RAM is `benchmark.pl`'s own figure, the `VmRSS` delta of a forked child running the call once, median of nine. The string row is there to show where the win is not: it is confined to numeric cells. - | Call | Time before | Time after | RAM before | RAM after | |---|---|---|---|---| | `drop_duplicates($hoa)` | 5.45 ms | 1.88 ms (2.9x) | 5.36 MB | 1.45 MB (3.7x) | | `merge`, inner join on an integer key | 7.78 ms | 5.62 ms (1.4x) | 7.41 MB | 6.41 MB (1.2x) | | `merge`, inner join on a double key | 10.28 ms | 5.91 ms (1.7x) | 7.12 MB | 6.42 MB (1.1x) | | `value_counts` on a double column | 2.98 ms | 1.66 ms (1.8x) | 1.98 MB | 1.55 MB (1.3x) | | `value_counts` on a string column | 0.215 ms | 0.220 ms | 0.69 MB | 0.71 MB | - `group_by` and `pivot_table` were left alone: `group_by` hands the cell SV straight to `hv_fetch_ent`, so perl does the stringification internally and reaching it means byte-level `hv_*` calls and a change to how UTF-8 keys are handled, and `pivot_table` is pure Perl. 0.294 2026-08-07 CDT - bug fixes: https://www.cpantesters.org/cpan/report/368ca238-73ee-1014-a03f-97f1b88bf904 - `binom_test` was cross-validated against R 4.6.1 `stats::binom.test` and SciPy 1.17.1 `scipy.stats.binomtest` using their own test suites rather than cases invented here: SciPy's `TestBinomTest`, R's `binom.test(c(800,10))` from `tests/reg-tests-2.R`, the `?binom.test` example, and an R-generated corpus of 383 p-values and 1560 Clopper-Pearson bounds. They are in `t/binom_test.R.scipy.t`. Two fixes came out of it, both in the incomplete beta that every tail and confidence bound goes through: - Its continued fraction stopped after a flat 500 terms, but it needs about 0.25 sqrt(a+b) of them once the shape parameters are large, so it was quietly cut short at big `n`: `binom_test(10079990, 21000000, p => 0.48)` returned 0.996781946606 where R and SciPy both give 0.9966892187965, i.e. wrong in the fourth decimal of a printed p-value. The cap now scales with sqrt(a+b), and the front factor moved off differenced `lgamma` onto the same saddle-point form `dbinom` already used here. Agreement with R over these cases went from 9.3e-5 to 3.3e-13 relative. - The Clopper-Pearson bounds are found by bisection, which stopped at an absolute width of 1e-15, so a bound far below 1 came back with only four correct digits: `binom_test(1, 1000000000, alternative => 'greater', conf.level => 0.999)` gave 1.00053299e-12 against R's 1.00050033e-12. The stopping rule is now relative to where the bracket sits, and such bounds now hold about 1e-15. - Both fixes also help `t_test`, `var_test` and `cor_test`, which use the same function. One limit remains, pinned by the tests rather than left to chance: the upper bound for a handful of successes in a billion trials still carries about 1e-9 of relative error, because the complement branch of the incomplete beta cannot resolve a tiny `x` past the spacing of `1-x`. 0.293 2026-08-06 CDT - Fixed quadmath error https://www.cpantesters.org/cpan/report/83bcd9a2-9123-11f1-aac1-f3cd035a6881 - `fisher_test` was cross-validated against R 4.6.1 `stats::fisher.test` and SciPy 1.17.1 `scipy.stats.fisher_exact` using their own test suites rather than cases invented here: SciPy's 84-case R-generated corpus (`scipy/stats/tests/data/fisher_exact_results_from_r.py`, four numbers per case over two confidence levels and all three alternatives), its `TestFisherExact`, R's regression suite (`tests/reg-tests-1{a,b,d,e}.R`: PR#644, PR#1662, PR#4688, PR#10558, PR#18336, PR#17671 and the "exact fisher.test" entry) and the `?fisher.test` examples. They are in `t/fisher_test.R.scipy.t`. Three fixes came out of it: - The 2x2 hypergeometric density was built by differencing `lgamma`, which costs the back half of a large table's p-value: at a margin of 8.4e7, `lgamma` is about 1.4e9, where a double's spacing is 2.4e-7, and exponentiating that turns into a relative error of the same size. SciPy's gh-3014 case came out right to only seven digits. The density is now assembled from Loader's saddle-point binomial, which is how R's own `dhyper` avoids this and which `binom_test` already had in the file; its three terms stay O(1) whatever the margins are. Worst-case agreement with R over the 84-case corpus went from 2.1e-12 to 5.2e-14 relative, and gh-3014 from 2.2e-07 to 1.5e-16. - The R x C enumeration charged only its leaves against its safety cap, so a table wide enough to spend the time in the interior of the tree neither finished nor stopped: R's PR#4688 table (4x3, N = 16442), whose whole point upstream is that `fisher.test` must fail rather than return `p = Inf`, ran for over five minutes here without doing either. Every node is now counted, and that table is declined in about a second. - The R x C enumeration now bounds each subtree before walking it. `lgamma(x+1)` is convex and `a!b! <= (a+b)!`, which together bracket the probability of every completion of a partial table; when the whole subtree falls inside the tail its mass is added in closed form (`N'! / (prod R_i! prod C_j!)`, from counting the remaining observations into rows two ways), and when it falls outside the subtree is dropped. The margins are also transposed and sorted first, so the fattest row and column are the ones the enumeration gets for free. R's Job Satisfaction 4x4 example went from 7.5s to 0.3s and PR#644's 19x2 from 1.0s to under 0.05s, and the 6x6 table of PR#18336 -- which segfaulted R before 4.2.0 and which R 4.6.1 still declines with `hash key 5e+09 > INT_MAX` -- is now computable at 0.6322160531, agreeing with R's own 2e6-replicate `simulate.p.value` fallback to within its sampling error. - Two behaviours that the two references disagree about are now pinned by tests rather than left to chance: a table with an empty row or column returns R's `p = 1` with an odds ratio of 0 and a CI of (0, Inf), not SciPy's NaN odds ratio; and a table with a single row or column is rejected as R rejects it, rather than returning SciPy's `p = 1`. 0.292 2026-08-05 CDT - fixed long-double bug https://www.cpantesters.org/cpan/report/506975f6-906a-11f1-8f30-a201c4f2440e - `power_t_test` was cross-validated against R 4.6.1 `power.t.test` and against `scipy.stats.nct` driven by `scipy.optimize.brentq`, over a grid of 288 cases covering all five solved-for parameters, all three types, both alternatives and `strict`. Three fixes came out of it: - The Simpson sum behind the noncentral *t* CDF put a fixed 30000-step grid on `u = w/(1+w)`, and the chi density it integrates defeats that at both ends. The density carries `w**(df-1)`, so unless `df` is a whole number some derivative of it is infinite at `w = 0` and Simpson's error bound does not hold: two good digits at `df = 1.2` with `sig_level = 1e-4`, five at `df = 1.2`, nine at `df = 1.8`. Substituting `w = z**m`, with `m` chosen so that `m*df - 1 >= 3`, restores the bounded derivatives and brings all of those to machine precision. It also puts the origin's contribution at zero, which subsumes a separate bug: the sum had been dropping its `u = 0` endpoint term, worth 7e-7 of absolute power at `df == 1`. `nu` is now also floored the way R floors it, per sample rather than in total. - The same density has standard deviation `1/sqrt(2*df)` and so narrows without bound, while the grid did not. Past `df` of about 1e7 the steps went clean over the peak: `power_t_test(n => 4e7, delta => 0)` returned 0.138 where the answer can only be `sig_level/2`, and a large-cohort `n` solved 9% low. Above `df` of 1e3 the steps now go on `w` across +/- 12 standard deviations of the mode, with the chi normalisation taken from Stirling's series to keep the peak height from cancelling away; and above 4e5, where those log terms cancel too hard for any grid to help, the Abramowitz & Stegun 26.7.10 asymptotic form takes over -- the same formula, at the same cut-off, that R's `pnt.c` uses. That is also 25 times quicker than integrating. - The power was formed as `1 - P(T <= t)`, which loses most of its digits to cancellation when the power is small. It is now integrated as the upper tail directly. - The four inverse solvers were plain bisection stopped at the bracket width, which capped `n`, `delta`, `sd` and `sig_level` at R's own four or five significant figures. They now use regula falsi with the Illinois correction against a relative tolerance, so they match machine-precision `brentq` roots to ~1e-13 in fewer evaluations than the bisection took. The `tol` default moved from `1.22e-4` to `1e-12` to match. - Nothing checked that the bracket held a root, so an unreachable target came back as a bracket endpoint wearing the requested power: solving for `sd` with `power => 0.01` returned `delta * 1e7`, and with a negative `delta` returned a negative standard deviation. Unreachable targets now croak and name the range searched. `sig_level` and `power` outside `[0, 1]`, an `n` below 2, a negative `sd`, and an unrecognised `type` or `alternative` are rejected as well -- `type => 'twosample'` used to be read silently as `'two.sample'`. - New test file `t/power_t_test.R.scipy.t` carries the cross-validated grid. 0.291 2026-08-04 CDT - POD formatting improvements [`lm`, `glm`] - Formula parsing and data reading are now shared between `lm` and `glm` too, so the two agree on what a formula means and on what a row is called. `lm` had the better parser and `glm` the better row naming; each now has both. - `lm` now names rows the way `glm` does — from a `row.names`, `_row`, `rownames` or `.rownames` column when the data has one, and 1-based integers otherwise. `lm` previously always used integers, so `fitted.values` and `residuals` came back keyed `1..n` for data whose rows had names, and did not match what `glm` or `predict` returned for the same data; the `predict` documentation already described the shared behaviour. A row-name column is a label rather than a measurement, so `y ~ .` now excludes it in both. - Design-matrix construction is now shared between `lm` and `glm`, and decides a categorical column's coding term by term using R's margin rule: the reference level is dropped when the term with that column removed is itself in the model. Three bugs fall out of that, all confirmed against R 4.6.1 and statsmodels 0.14.6. - Bug fixes: - Four in `glm`, from the parser it now shares with `lm`. Three of them ended the same way: a term that names no column evaluates to `NaN` for every row, every row is dropped as incomplete, and the fit dies with `0 degrees of freedom (too many NAs or parameters > observations)` — never mentioning the formula. - **`glm` truncated a formula at 511 characters.** It copied the formula into a fixed `char[512]`, so a model with enough predictors to overrun that lost the tail. The buffer now grows with the formula. - **`glm` did not understand `.`.** It parsed the formula before reading the data, so there were no column names to expand `.` into and the term stayed a literal `.`. Formula splitting now happens first and term expansion after the data is read, so `y ~ .` works in both. - **`glm` did not understand `+ 0` or a leading `0 +`.** Only `- 1` suppressed the intercept; the other two spellings R accepts left a term named `0`. All three now work in both, as do `+ 1` and a leading `1 +`. - **`glm` read the `-1` inside `I(...)` as intercept suppression.** It searched the whole right-hand side for the substring, so `y ~ I(x-1)` silently became `y ~ I(x) - 1`: a different model, fitted without complaint. The scan now steps over `I(...)`, leaving the term alone. `I()` still supports only `^power`, so that formula is an error in both rather than a wrong answer in one. - And the three that fall out of the shared design matrix: - **A categorical column in a model with no intercept lost a level.** With no intercept there is no baseline for a reference level to be measured against, so R codes the factor in full — one column per group, each coefficient that group's own mean. Both functions dropped the reference level anyway, so `len ~ supp - 1` fitted `len ~ suppVC - 1`: a model forcing every observation at the reference level to a fitted value of 0. On R's `ToothGrowth` that meant a residual sum of squares of 16056 against R's 3247, and an R² of 0.35 against 0.87. Where two categorical main effects appear with no intercept only the first is coded in full, as in R, since coding both would be rank deficient. - **An interaction involving a categorical column could not be built.** The interaction was looked up as a single column literally named `dose:supp`; finding none, it evaluated to `NaN` for every row, every row was dropped as incomplete, and the fit died with `0 degrees of freedom (too many NAs or parameters > observations)`. Interactions now expand to the product of their components' indicator columns, so `len ~ dose * supp` gives `dose`, `suppVC` and `dose:suppVC`. `predict` already understood such coefficient names; now they can be produced. - **`a*b*c` expanded only its first `*`.** Crossing is associative, so `y ~ a * b * c` now yields every non-empty subset (`a`, `b`, `c`, `a:b`, `a:c`, `b:c`, `a:b:c`), ordered by degree as R's `terms()` orders them. Previously the chunk was split once, producing the unusable terms `b*c` and `a:b*c`, and the fit died the same way as above. Crossing more than 16 columns now croaks rather than expanding to 2^n terms. - **`predict` scored reference-level rows as if the term were absent.** It registered factor dummies from `levels[1..]` only, on the assumption that a reference level never has a coefficient — true for a factor coded by contrasts, but not for one coded in full. Every row at the reference level of a no-intercept model therefore came back 0. - **`glm` halved its IRLS step whenever the deviance rose, costing iterations and accuracy in the standard errors.** R truncates a step only when the deviance comes out non-finite; a deviance that merely increases is not divergence. The standard IRLS start puts `mu` at `y + 0.1`, essentially on the data, so the initial deviance is near zero and the first real step almost always raises it — on the nine-point poisson fit in `t/glm.t`, from 0.016 to 1.54. That was read as divergence and the step was halved ten times over, turning R's four iterations into seven. - The extra iterations reached the same coefficients, so the symptom appeared only in the standard errors. They are built from the information matrix of the *penultimate* iterate — in R because `summary.glm` inverts the QR that `glm.fit` kept from its last weighted least squares call, and here because the IRLS sweep leaves that inverse in place — so stopping on a different iteration than R means reporting a different matrix. Poisson standard errors were 5e-8 to 2e-5 away from R's while the coefficients agreed to twelve digits; they now agree to about 1e-14. Binomial standard errors were up to 6e-7 out and now agree to 2e-14, except on a near-separable fit, where the `varmu` floor of 1e-10 (a guard against dividing by an underflowed variance) accounts for the remaining difference — 1.9e-9 on `am ~ wt * hp`, where three of 32 fitted probabilities are within 1e-12 of 0 or 1 and R itself warns. Gaussian fits are unaffected: their weights are all 1, so the matrix is `X'X` either way. - The same condition had its `isfinite` test on the accepting side, so a genuinely divergent step producing a non-finite deviance was kept rather than truncated. That is now the one case that does trigger halving. - **The negative-binomial theta alternation stopped early and started from the wrong place.** `MASS::glm.nb` does not simply maximise over theta; it alternates between an IRLS fit at the current theta and a fresh ML estimate of theta at the current fitted means, and which fit it lands on depends on the schedule. Four details of that schedule were wrong here, and all four are now reproduced: - The alternation stopped on a relative test of the log-likelihood alone, `|dll| < 1e-7 * (|ll| + 0.1)`. `glm.nb` requires `(|dLm| / d1 + |dtheta|) < 1e-8` with `d1 = sqrt(2 * max(1, df.residual))` taken from its Poisson pass — theta itself has to have settled, not just the log-likelihood. The old test was satisfied roughly 2e-5 of log-likelihood early, which left theta 8e-7 out and dragged the coefficients 8e-6 with it. - The first pass now runs as a genuine **Poisson** fit, as `glm.nb`'s does, rather than a negative-binomial fit at a large stand-in theta. That pass supplies both the first theta and the `d1` above. - Later passes are **warm started** from the previous pass's means (`etastart = log(mu)`), so they converge to the fit `glm.nb` reaches rather than to the same optimum approached from a cold start. - Theta is re-estimated at the means each pass **started** from, not the ones it produced: `glm.nb` calls `theta.ml(Y, mu)` and only then reassigns `mu <- fit$fitted.values`. `theta.ml` itself now also uses MASS's own stopping rule, an absolute Newton-step tolerance of `.Machine$double.eps^0.25`. - Across eighteen fits spanning dispersion from theta 0.41 to theta 69000, theta now agrees with `glm.nb` to 3.4e-9, coefficients to 5.8e-9, standard errors to 8.4e-10 and deviance to 1.3e-9 — previously 8e-7, 8e-6, 3e-6 and 6e-7. The one exception is genuinely near-Poisson data, where theta is not identified at all (its own standard error exceeds the estimate, and `glm.nb`'s `theta.ml` reports "iteration limit reached"); theta there agrees only to about 4e-6 relative, while the coefficients still agree to 1.6e-10. - A separate consequence: a negative-binomial fit with theta supplied was starting from the Poisson `mustart` of `y + 0.1`, where R's `negative.binomial()$initialize` sets `y + (y == 0)/6`. Different starting values walk different iterates, and since the standard errors come from the penultimate one, that showed as standard errors 6e-7 from R's while the coefficients agreed to 1e-9. Such fits now match R to 2e-15. - Note on that comparison: standard errors for a negative-binomial fit hold the dispersion at 1, which is what `glm.nb` and `summary.negbin` do. R's `summary.glm`, handed a `negative.binomial` family directly, instead *estimates* the dispersion and prints standard errors scaled by its square root — 1.0839 on one of the test data sets, so about 4% larger. Compare against `summary(fit, dispersion = 1)` to see the values this module reports. 0.29 2026-08-03 CDT [t_test] - `t_test` was cross-checked against R's `stats::t.test` and `scipy.stats` case by case, including the cases their own suites pin: R's regression tests (`reg-tests-1a.R`, "t.test with one group of size one") and scipy's `TestTTest_1samp`, `TestTTest_ind.test_special_cases`, `test_ttest_rel_ci_1d`, `test_1samp_ci_1d` and `test_pvalue_ci`. On 2000 randomised comparisons against R — all four modes, all three alternatives, random `mu` and `conf.level`, sample sizes 2 to 40 and data scales spanning 1e-4 to 1e4 — the statistic and the degrees of freedom agree to 2e-11 and the p-value to 3e-9, holding to eight digits even where the p-value is subnormal (5e-310). What the comparison did turn up was seven ways a call could come back wrong rather than loud, all of them now fixed and covered by `t/t_test.t`. - **`undef` was coerced to 0 instead of being dropped.** This is the one worth re-running results over. `t_test` did not filter missing values, so a column with gaps in it was tested with every gap counted as a zero: R gives `t.test(c(1,2,NA,4,5))` a `t` of 3.286 on 3 degrees of freedom, and `t_test` answered 2.588 on 4. No error, no warning, and an answer close enough to the real one to look right. `undef` and `NaN` are now dropped the way R drops `NA`, per-vector for a one-sample or unpaired test, and on complete cases when `paired` so a half-missing pair goes whole rather than contributing a difference against zero. - **A `y` of fewer than two observations returned a silent `NaN`.** `var_y` divided by `ny - 1`, so `t_test(\@x, [$one_value])` propagated `0/0` into the statistic, the p-value and both interval bounds without raising. The two thresholds R uses are now both in place: a Welch test needs a variance from each side and refuses without one, while a pooled test tolerates a side of one observation, since that side contributes no sum of squares. That second case is what R's own regression suite pins — `t.test(y=x[1], x=x[-1], var.equal=TRUE)` is a well-defined test with 8 degrees of freedom, and `t_test` now answers it instead of returning `NaN` in one direction and croaking in the other. An empty `y` is caught by the same check. - **`alternative` was never validated.** The p-value helper fell through to two-sided for any string it did not recognise, so a typo — `'gerater'` — ran a different test than the caller asked for and reported nothing. It is now checked the way R's `match.arg` checks it. `scipy`'s `"two-sided"` spelling is unambiguous, so it is accepted rather than rejected. - **A one-sided interval was wrong when `conf.level < 0.5`.** That case needs a negative t quantile, and `qt_tail` searched upward from zero only, so it returned roughly zero and collapsed the bound onto `mu`: R puts the upper bound of `t.test(1:10, mu=5, conf.level=0.3, alternative="less")` at 4.9797 where `t_test` reported 5.0000000036. `qt_tail` now reduces by symmetry first, so the root it brackets is always positive. - **`qt_tail` silently saturated at 1e6.** Past that its doubling loop gave up and returned the ceiling, so `conf.level` of 0.99999999 and 0.9999999999 came back with the *identical* interval, ±1048576, against R's ±6.4e7 and ±6.4e9. The ceiling is gone; the loop now runs until `t * t` would overflow. - **Interval accuracy no longer depends on the data's scale.** `qt_tail` bisected to an absolute 1e-8 on the quantile, which is 1e-8 × `std.err` on the interval — fine for data around 1, an error of 2 units for data around 1e9. It now bisects to adjacent doubles. Worst interval error across the 2000 randomised cases went from 2.1e-4 to 5.3e-11 relative. At extreme `conf.level` this makes `t_test` the more accurate of the two: `t.test` asks for `qt(1 - alpha/2, df)`, and representing a 5e-9 tail as the double `1 - 5e-9` costs eight significant figures of it, so R's own interval for `conf.level=0.99999999` is off by 0.7 in the eighth digit. Working in the upper tail throughout agrees with R's `qt(alpha/2, df, lower.tail=FALSE)` to 15 digits. - **"Essentially constant" was an absolute test.** Only an exactly-zero variance was rejected, so a spread below what a double can resolve at the data's own magnitude was reported as a finding: four values around 1e10 differing by 1e-5 gave `t` = 4e15 and a p-value of 3e-47. The comparison is now relative, as R's is. The exactly-zero case, where R returns `NaN`, raises here instead. - **A defined non-array `y` was ignored.** `t_test(\@x, y => 5)` quietly ran a one-sample test. It now raises. An explicit `undef` still means absent, as R's `y = NULL` does. - `qt_tail` is shared with `power_t_test`, which gains the same precision; it is only ever called there with a tail below 0.5, so nothing about its behaviour changes. `t_test` remains allocation-free — the missing-value filtering happens inside the same single Welford pass that was already there, and the result hash is built after the last error check rather than before the first. [write_table: `tex.longtable.head`] - A `longtable` freezes only the header sitting inside `\endfirsthead` / `\endhead`, and `tex.longtable` never wrote those blocks — its header was an ordinary first body row, leaving the frozen one to be hand-written by the caller. That header then had no link to `col.names`: reorder the columns and the labels at the top of every page keep the old order while the data below them moves, and the generated header appears again as a duplicate first row. - `tex.longtable.head` generates the repeat machinery from the table's own header record, so it cannot drift. A true-but-numeric value emits `\endfirsthead`/`\endhead`/`\endfoot` with no continuation caption; any other true value is the caption used on pages after the first, written verbatim. Implies `tex.longtable`. `tex.longtable` on its own is unchanged. - The wrapper keeps one static token, the `\hline` closing its `\caption` line, because a leading `\hline` in an `\input`ed file is a `Misplaced \noalign` error — TeX has already begun the row by the time it expands the `\input`. [skew, kurtosis] - Two new XS functions describing the shape of a sample beyond its spread: `skew` for the third central moment and `kurtosis` for the fourth. Both take arguments the way `sd` and `var` do — numbers, array references or a mixture, flattened into one sample — and both also accept `x => \@data` and `type => 1|2|3`. - `type` selects among the three sample conventions, which disagree noticeably on small samples. The default is `type => 2`: `G1` and `G2`, the estimators unbiased for a normal sample, as reported by SAS, SPSS, Stata, Excel's `SKEW()` and `KURT()`, and `scipy.stats` with `bias => FALSE`. `type => 1` is the plain moment ratio (`moments::skewness`) and `type => 3` is `b1`/`b2` (`e1071::skewness`'s own default). All three, for both functions, agree with R to about 1e-15. `kurtosis` returns *excess* kurtosis — 3 is already subtracted, so a normal sample sits near 0. - One pass, no allocation: the third and fourth central moments accumulate through Welford's recurrence extended to higher moments (Terriberry) rather than the textbook expansion in raw moments. That expansion is not usable on real data — for a column of values around 1e7, a lab value in the wrong units or a timestamp, `sum(x**3)/n` is about 1e21 while the third central moment is single digits, so every significant figure cancels away. - A constant sample croaks rather than returning a silent `NaN` from `0/0`, and a `type` whose denominator the sample is too small for (`type => 2` needs `n >= 3` for `skew` and `n >= 4` for `kurtosis`) says which. - Both read tied arrays. `av_fetch` on a tied array returns a deferred `PVLV` rather than the value, and `SvOK` on one of those is false until its get-magic has run, so without an `SvGETMAGIC` every element of a tied array looks undefined. [median] - `median` now reads tied arrays too. It already had a separate `av_fetch` path for them — a tied array keeps nothing in `AvARRAY`, so the fast path would read off a null pointer — but that path was missing the `SvGETMAGIC` described above, so it rejected every tied array as undefined instead of computing the answer. `mean`, `sd`, `var`, `sum`, `min` and `max` still reject tied arrays. They have no `AvARRAY` fast path to guard, so they croak rather than crash, and the same one-line fix would make each of them work. [oneway_test] - `oneway_test` was cross-checked case by case against R's `stats::oneway.test` (both branches), R's `anova(aov())` for the `Sum Sq` / `Mean Sq` columns, `statsmodels.stats.oneway.anova_oneway(use_var="unequal")` and `scipy.stats.f_oneway`. The 37 data sets are R's own built-ins — `chickwts`, `InsectSprays`, `PlantGrowth`, `iris`, `ToothGrowth`, `mtcars`, `warpbreaks`, `sleep`, `airquality`, `CO2`, `esoph`, `OrchardSprays`, `faithful`, `quakes` — plus hand-built numerical edge cases. The statistic and both degrees of freedom already matched R everywhere; what the comparison turned up was four ways a call could come back wrong rather than loud, all now fixed and covered by `t/oneway_test.R.scipy.t`. Statistic, degrees of freedom and p-value now agree with R to 1.3e-12 relative error across all 37, and on 2000 randomised comparisons against R — both branches, 2 to 8 groups, sizes 2 to 40, deliberately heteroscedastic, data scales 1e-4 to 1e4 — the statistic and the degrees of freedom agree to 1e-12 and the p-value to 8e-11, the worst of those being a p-value of 2.4e-66. - **Every p-value below about 1e-16 was returned as a flat 0.** `Pr(>F)` was built as `1 - pf(F, df1, df2)`, and 1 minus something that close to 1 has no bits left to carry the answer: `faithful` split at `waiting > 70` should give `1.2099104551915e-76` under Welch and `5.50783574504386e-103` pooled, and `oneway_test` reported `0` for both. Anything from about 1e-9 downward was losing relative precision the same way, quietly — `ToothGrowth` by dose came back as `9.99200722162641e-16` against R's `9.53272701169993e-16`, off by 4.6% with nothing to indicate it. The p-value is now evaluated in the upper tail directly, via the beta symmetry `1 - I_x(a, b) = I_{1-x}(b, a)`, so no subtraction from 1 happens at any point, and the range down to the smallest representable double is reported at full precision. - **An `F` of `Inf` produced a p-value of `NaN` instead of 0.** When every group is constant but their means differ, the within-group sum of squares is 0 and `F` is legitimately infinite; R reports `p = 0`. `pf` formed `df1*f/(df1*f + df2)`, which is `Inf/Inf` — a `NaN` that propagated straight into `Pr(>F)`. `Inf` and `NaN` are now handled explicitly, matching R's `p = 0` and `p = NaN` respectively. - **A `NaN` Welch denominator df was reported as 1e300.** A group with zero variance gets an infinite Welch weight, which makes R's `tmp` term `NaN` and its denominator df `NaN` with it. `oneway_test` had a `(tmp > 0.0)` guard that a `NaN` fails, so it substituted a magic `1e300` — a number that reads as a real, very large degrees of freedom and would be believed as one. The guard is gone; `Residuals`/`Df` and `Residuals`/`Mean Sq` are `NaN` there, as in R. - **`formula` mode read `undef` and non-numeric response cells as 0.0.** The hash and array-of-arrays shapes were already fixed to die on these (and pinned by `t/oneway_test.bugs.t`), but the formula path has its own fill loop and was missed, so oneway_test({ y => [1, 2, 3, undef, 5, 6], lab => [qw(a a a b b b)] }, formula => 'y ~ lab'); - silently tested group `b` as `(0, 5, 6)` — a mean of 3.67 instead of 5.5 — and returned an `F` of 0.735 with no complaint. All three input shapes now enforce the documented contract identically. - Two places where `oneway_test` is the more accurate side and the reference is not, now documented rather than treated as disagreements: the sums of squares are accumulated two-pass, so on two groups near 1e8 `Residuals`/`Sum Sq` is exactly `10` where R's QR-based `anova(aov())` gives `10.0000000521067`; and where the exact between-group sum of squares is 0, `oneway_test` returns 0 rather than R's 1e-30-scale residue. 0.281 2026-08-03 CDT - `median` (LikeR.xs) — the same answers in about an eighth of the time. On the `benchmark.pl` case (10,000 normals in one array ref) a call went from 0.83 ms to 0.097 ms, which puts it ahead of the two implementations it was behind: `numpy.median` at 0.105 ms and R's `median` at 0.196 ms, measured on the same machine. Small samples — a per-group median under `agg` or `group_by`, which is where most calls to it come from — went from 647 ns to 251 ns. - A median is the middle one or two values, so most of the sort the function used to do was wasted work: `qsort` orders all n elements at a cost of n log n comparisons, every one an indirect call through a function pointer the compiler cannot see into, to answer a question that depends on one or two of them. Those values are now selected instead, and the sample is walked once rather than twice. - The selection is introselect, the same shape numpy's `partition` uses: quickselect with a median-of-three pivot, an insertion sort once a range is small, and a heapsort fallback past a depth limit, so an input crafted to defeat the pivot choice degrades to O(n log n) rather than O(n²). The awkward data people actually have comes out faster than random data rather than slower — sorted, reversed, all-equal and organ-pipe samples of 100,000 values each take about 0.25 ms against 0.90 ms for random ones. For an even count the lower of the middle pair is the largest value left below the upper one, which a scan of that side finds without a second selection. - The counting pass is gone. It walked every element through `av_fetch` before any arithmetic, only to size the buffer; the array lengths give the same count, exactly, because an undef anywhere still dies. - The pass that remains reads cells through `AvARRAY` instead of `av_fetch`, with tied arrays kept on the `av_fetch` path, since only it sees their values. - A sample of 256 values or fewer is copied to the C stack rather than the heap, so the common small call no longer pays for a malloc and free at all: 10,000 of them now grow RSS by nothing. Larger samples still copy n values, which is what leaves the caller's array in its original order — the selection reorders whatever it works on, and `t/median.t` checks that the input comes back untouched. - Error messages that carry an index or a count were unreadable on older perls, in nineteen places across LikeR.xs, and now are not. `croak` runs perl's own formatter rather than the C library's, and that formatter does not understand C99's `z` length modifier: it printed the conversion literally, so `median(1, undef)` on perl 5.10 or 5.12 said `undefined value at argument index %zu` instead of naming the argument that was undefined — the one thing the message existed to say. `min`, `max`, `mode`, `sum`, `sd`, `var`, `median`, `mcnemar_test`, `friedman_test`, `hoa2hoh` and `oneway_test` were all affected. They now use `UVuf`, as the rest of the file already did. The `snprintf` calls elsewhere in the file are unaffected and unchanged: those do go to the C library, where `%zu` means what it says. - New `t/croak.messages.t` covers every one of those messages: each is triggered, checked for the number it should name, and then swept for any conversion left unexpanded, which is what will catch the next one written with `%zu`. Run against the code as it stood before this change, thirty-two of its assertions fail on perl 5.10. - New `t/median.t`: every length from 1 to 25 and from 254 to 258 — either side of the insertion-sort cutoff and of the point where the buffer moves off the stack — across sorted, reversed, all-equal, two-valued, duplicate-heavy, organ-pipe and median-of-three-killer samples, each checked against a plain Perl sort, together with the error messages and `Test::LeakTrace` over the stack, heap, mixed-argument and croak paths. - fix for threaded Perls https://www.cpantesters.org/cpan/report/2dbacf8f-7138-1014-a1ab-f0f91cf3b922 0.28 2026-08-02 CDT - `p_adjust` (LikeR.xs) now takes a data frame as well as a flat list of p-values, and hands the corrected values back in the shape they arrived in. An AoA, AoH, HoA or HoH goes in and a new frame of the same kind comes out, with the same rows, columns and row labels; the input is left alone. Everything the flat form did is unchanged — an arrayref of p-values still returns a list, in order, with the same numbers. - `columns => 'p_value'` (or an arrayref of names, or 0-based positions for an AoA) says which columns hold p-values, and copies the rest of the frame through untouched, so a results table with a `gene` column no longer has to be taken apart and put back together around the call. Without `columns` every cell is treated as a p-value, which is right for a frame that is nothing but p-values; a label column in one dies with a message naming the offending value and pointing at `columns`, rather than correcting a string coerced to zero. - All the p-values in the frame are corrected as one family, whichever shape they came in, so the family size is the number of p-value cells. - The method still reads positionally and may now also be given as `method => ...`. `none`, which the function has always accepted, is now documented along with the rest. - Cells are visited in a fixed order — by row and then column name, or column name and then row for a HoA — so tied p-values break the same way on every run instead of following hash iteration order. - `drop_duplicates`, `filter`, `t_test`, `vals`: speed/RAM improvements - Incompatible: the `'?'` / `'h'` argument added in 0.27 is gone (lib/Stats/LikeR.pm). `agg('h')`, `read_table('?')` and the fifty-odd other pure-Perl functions that took it no longer print help and die — they treat the string as data, the way the XS functions always have. `h('agg')`, `h(*agg)` and `h(\&agg)` are unchanged and remain the way to ask, for every function in the distribution. - It was a help route that only half the module had, so what a lone `'h'` meant depended on whether the callee happened to be written in XS or in Perl, and a column, file or option value really named `'h'` needed `$Stats::LikeR::HELP = 0` to get through. That variable is gone too; nothing reads its arguments for a help flag any more. - `bedroc` still prints its own short XS usage summary for `bedroc('h' | 'H' | '?')`, which is hand-written and predates all of this. - `merge` (LikeR.xs) — same joins, a third of the time and a fifth of the memory. Nothing about the result changes: every join type, shape combination and edge case produces exactly what it did before, and `t/merge.t` now checks all six input/output paths against a plain-Perl reference join over a randomized corpus. - The old implementation transposed both frames into arrays of row hashes, joined those, and transposed the result back. A 10,000-row HoA joined to itself therefore built 20,000 throwaway row hashes and copied every cell three times before returning. It now reads each frame where it lies — a HoA column by column, an AoH/HoH row by row — and writes the result straight into the shape being returned, so the only cells copied are the ones the caller keeps. - The right frame's index is a hash of row numbers chained through a flat array, rather than an array-ref of index scalars per distinct key, and one reused buffer builds every join key instead of one scalar per row. - Column names are resolved to their column (HoA) or interned once as shared hash keys (AoH/HoH) before the join starts, so the per-row work is a lookup rather than a lookup and a rehash. - Measured on the `benchmark.pl` case (two 10,000-row frames, six columns, inner join on `id`): 0.052 s and 41.4 MB before, 0.017 s and 7.3 MB after. An outer join of the same frames went from 0.113 s to 0.008 s. - `write_table` (LikeR.xs) — two changes, one of them incompatible. - Every format now prints the coloured `wrote ` confirmation line, not just LaTeX and `.xlsx`. Delimited output (csv/tsv) was silent before. The line is identical in all cases: the file name in black on cyan, with the SGR codes inline so there is still no `Term::ANSIColor` dependency. Nothing is announced when nothing is written. - **Incompatible:** `row.names` now defaults to **off** in every format. It previously defaulted **on** everywhere, following R's `write.table`, which meant a call that said nothing about row names got a label column and a leading empty header cell (`,gene,n`) it had not asked for. Pass `row.names => 1` for the old behaviour; `row.names => 'col'` is unchanged. - New `h2aoh` and `aoh2h` (lib/Stats/LikeR.pm), which add the flat hash to the shapes the conversion family understands. A plain hash is a two-column table folded shut, and until now nothing would unfold it: `value_counts` hands one back, and no frame function would take it. - `h2aoh(\%h, var_name => .., value_name => ..)` unfolds a flat hash into a two-column AoH, one row per pair, under column names the caller picks. `sort => 'key' | 'value' | 'none'` fixes the row order, which hash iteration otherwise leaves to chance; `'value'` is biggest-first for numbers, so `value_counts` output comes out the way pandas' `Series.value_counts()` orders it. - `aoh2h` folds a two-column AoH back down, with `duplicates => 'die' | 'first' | 'last'` deciding what a repeated key means. The two are exact inverses under their defaults. - The column options are named `var_name` / `value_name` after `melt`, which emits the same two columns. R spells this pair `tibble::enframe()` / `deframe()`; pandas spells it `pd.Series(d).reset_index()` and `Series.to_dict()`. 0.27 2026-07-26 CDT - New `h` function: `h('agg')`, `h(*agg)` or `h(\&agg)` prints that function's section of this document and returns, in the spirit of R's `?function`. `h()` lists every documented function. It covers the XS functions as well as the Perl ones, because it looks the name up in the module's POD instead of reading an argument list — see [Getting help](#getting-help). - The pure Perl functions also accept `'?'` or `'h'` in place of their arguments, which prints the same text and then dies. `$Stats::LikeR::HELP = 0` switches that off for code that has to pass a column or file really named `'h'`. - `qcut`'s hand-written usage message was replaced by its section of this document; `qcut('h')` and `qcut('?')` still die, but `qcut('H')` no longer means help. - speed improvements in calculation of Kendall tau and p-value. Improvement of writing xlsx files that won't show in time, but pure waste was removed. - Addition of `auc`, `auroc`, `cmh_test`, `epi_2x2`, `roc` functions - `prcomp` now accepts AoH input - glm extended (LikeR.xs) - family => 'poisson' (log link) and family => 'negbin' — negative-binomial θ estimated by ML via a MASS::glm.nb-style outer loop, or fixed with theta =>. Matched R to ~1e-8 (coefs, deviance, null-dev, AIC, SE, θ); exact Poisson limit when data aren't over-dispersed. - Every non-gaussian family now returns exp (odds/rate/incidence-rate ratios + conf.low/conf.high), link-scale conf.int, conf.level, and theta (negbin). Count families report z-statistics. OR/CI matched R's confint.default exactly. - New XS tests (all matched R exactly) - prop_test — 1/2/k-sample proportions (Yates, Wilson & Wald-diff CIs) - mcnemar_test — matrix or paired vectors; continuity correction; exact => 1 binomial - friedman_test — repeated-measures rank test, tie-corrected - dunn_test — post-Kruskal pairwise, 7 adjustment methods - New Perl functions (lib/Stats/LikeR.pm, matched base-R references) - Effect sizes: cohen_d (+Hedges g, CI), smd, cramers_v (+Bergsma bias-corrected), eta_squared (η²/partial/ω²) - vif, hosmer_lemeshow (matches hoslem.test) - age_standardize — direct standardization + Fay–Feuer gamma CI (matches epitools::ageadjust.direct) 0.26 2026-07-20 CDT - https://www.cpantesters.org/cpan/report/fc7d01a0-83f4-11f1-b543-8a9ac547de9a Fixed a long-double issue 0.25 2026-07-19 CDT - https://www.cpantesters.org/cpan/report/3376f80e-83bf-11f1-a5f3-44496e8775ea - Fixed a use-after-free in `fisher_test` on the hash (HoH) input path: the "row is missing column key" error freed its scratch arrays and then read the key strings back out of them to build the croak message. This was harmless on glibc but crashed (`SIGBUS`) under stricter allocators such as FreeBSD's, failing `t/fisher_test.t` on CPAN smokers. The key pointers are now captured before the arrays are freed. 0.24 2026-07-19 CDT - `interpolate`'s numeric core moved from pure Perl to XS (`_interp_column_xs`): ~5× faster for `linear` on large columns, ~11× for `pchip`, and ~50× for the spline methods whose dense solve dominates. Results are unchanged (bit-for-bit versus the former Perl kernels). - `Ronly` now accepts one or more array references (like `Lonly`), returning the values found only in the last reference; the two-argument form is unchanged, and `Ronly(@refs)` equals `Lonly(reverse @refs)`. - `interpolate` gains full `pandas.DataFrame.interpolate` method parity: `nearest`, `zero`, `slinear`, `pad`/`ffill`, `bfill`/`backfill`, `quadratic`, `cubic`, `cubicspline`, `pchip`, `akima`, `barycentric`, `krogh`, `polynomial`, `spline`, and `index`/`values`/`time`, plus an `x` argument for custom abscissae and an `order` argument. Matched to pandas/scipy within 1e-6. - `t/transpose.t` no longer loads `Devel::Confess` in its leak tests: its `$SIG{__DIE__}` stack-trace objects landed in `$@` and were reported as leaks by `Test::LeakTrace` on the croak paths under older perls (e.g. 5.12.3). The die-path leak checks now also clear `$@` so the exception object cannot be miscounted. - `cfilter` simplification, use of `qr///` filtering on columns - `summary` output now looks more like `view`, and accepts HoH - `fisher_test` can compute larger tables than just 2x2 - `read_table` reads xlsx files significantly faster and with less RAM. - Addition of `bfill`, `drop_duplicates`, `ffill`, `melt`, and `pivot_table` - Original `Lonly` code removed, as it was a special case of `get_unique`, and `get_unique` was re-named to `Lonly`. - Removal of `Devel::Confess` from testing and dependencies. 0.23 2026-07-10 CDT - `rename_cols` takes HoH as input - `write_table` prints row names as first column; writes longtable with comments - `assign` gains `map_cell { ... }` for in-place per-cell column edits [assign] - `assign` now accepts a third kind of column value, `map_cell { ... }`, for editing an existing column in place — no "copy, substitute, return" boilerplate and no dependence on `s///r` (unavailable on the older perls this module supports). - Inside a `map_cell` block, `$_` is the **named column's current cell** (not the whole row), the block's return value is **ignored**, and the modified `$_` is stored back: `assign($df, 'Res.' => map_cell { s/^[A-Z]:// })`. - The row is still available as `$_[0]` (sibling columns), the index as `$_[1]`, and the row key as `$_[2]` (HoH only). - **Undef/missing cells pass through untouched** (undef in → undef out): the block is skipped for them, so `s///` never warns on an uninitialized value. - Supported on all three shapes; for HoA the target column must already exist. A plain `sub { ... }` is unchanged, so `map_cell` is purely additive. - `map_cell` is exported alongside `assign`. - Tests: `assign.t` (AoH + HoA) and `assign.HoH.t` gained `map_cell` coverage — in-place `s///`, `$_[0]`/`$_[1]`/`$_[2]` context, new-column-from-undef, the missing-HoA-column death path, and `no_leaks_ok` guards. Verified building and passing the full suite on perl 5.10.1, 5.12.5 (long-double), and 5.42.2. [`group_by`] - Fixed group_by to honor all filter hashrefs (option 1) - Root cause: the XS captured only ST(3), so every filter hashref after the first was silently dropped — including the README's documented multi-hashref form. - Change (LikeR.xs): - Removed the single-ST(3) capture and the filter_hv PREINIT var. - Added a FOR_EACH_FILTER(body) macro that walks the arg stack from ST(3) to ST(items-1), iterating every { column => sub } pair and ANDing them together. It iterates the stack directly rather than heap-collecting the hashrefs, so a croaking filter sub still can't leak anything (verified). Non-hashref args are skipped. - Rewrote the filter loop in all three branches (AoH / HoA / HoH) to use the macro, keeping each branch's own value-fetch logic. - One build wrinkle worth noting: xsubpp parses every non-# line in the inter-XSUB region as a candidate function signature, so a /* ... */ comment there breaks the build (it tried to parse column => sub / (ST(3)..) as a signature). I moved the macro's documentation into the XSUB body (real C) and left the macro comment-free, matching the existing EVAL_FILTER style. - Tests (t/group_by.HoH.filter.t, 17 assertions): - HoH single-column filter, AND filter (both the one-hashref and separate-hashref forms now give identical results), no-match → empty hash, missing/undef target excluded despite passing the filter, and no_leaks_ok - Mentioning a non-existent column is now fatal. 0.22 2026-07-07 CDT - returned `Devel::Confess` to required dependencies to fix for CPAN testers. 0.21 2026-07-07 CDT - Better warning message for undefined data for `aoh2hoh`, `assign`, `dropna` - addition of `agg`, `concat`, `drop_cols`, `rank`, `rename_cols`, `select_cols` functions - Improving Kwalitee (sic): added `[PodWeaver]` to dist.ini; as well as `Changes` file [`assign`] - `assign` now accepts two kinds of column value, so a function that already returns a whole column (like `rank`) drops in without wrapping. - **Per-row coderef** (unchanged): called once per row, `$_` is the row, and the single scalar it returns is the cell. A single arrayref return is still stored *as the cell*, so arrayref-valued columns keep working. - **Whole-column coderef** (new): if the coderef returns a *list* of more than one value, that whole list becomes the column, laid down positionally. This is what makes `'ΔG rank' => sub { rank( vals($df, 'dG_kcal_mol') ) }` work directly — no `[ ... ]` needed. - **Arrayref value** (new): a ready-made column, e.g. `col => [ rank(...) ]`, copied into the frame. - The coderef is probed once (row 0 for AoH/HoH, the first synthesized view for HoA) to decide per-row vs whole-column, so per-row code is never run twice on row 0. Every column value is length-checked against the row count and a mismatch dies. HoH is now a supported, documented shape alongside AoH and HoA; whole-column and arrayref values align to sorted key order. - Tests: `assign.t` (AoH + HoA) and `assign_HoH.t` were expanded to cover every shape × value-kind combination — per-row scalar, whole-column list, arrayref value, single-arrayref-as-cell, `rank()` integration, chaining, `$_[1]` index, `$_[2]` row key (HoH), overwrite, ragged HoA columns, empty frames, length-mismatch and bad-value / odd-arg / non-hash-row death paths, and `no_leaks_ok` guards on the new whole-column and arrayref paths. [`read_table`] - Fixed handling of commented-out header lines and made filter columns referenceable by the name as it appears in the file. - **Commented-out header recovery.** `_parse_csv_file` treats a line whose comment marker is followed by whitespace (e.g. `# PDBscore`) as a comment and drops it, so a header written that way never reached the callback and the first *data* row was silently mistaken for the header. `read_table` now recovers it: the first physical line, if it is `marker + whitespace` and splits into two or more fields, is held as a candidate header and confirmed only when its field count matches the first data row. If the counts disagree the candidate was an ordinary leading comment and is discarded, so a prose comment that happens to contain the separator (e.g. `# note, see README`) is never mistaken for a header. A marker hugging its text (`#id,val`) is delivered by the parser and un-commented in the callback as before. The marker and any following whitespace are stripped, so `# PDB` is stored as the clean name `PDB`. - **Filter columns may be named as written in the file.** Filter keys are matched against the header by exact name first, then retried with the leading comment marker (and surrounding whitespace) stripped, so a commented-header column resolves whether it is referenced as `# PDB` or by its clean name `PDB`: read_table( 'regression_rank.tabular.tsv', filter => { '# PDB' => sub { $_ == 2 } }, ); - **Clearer "column not found" error.** The failure now names the file and lists the actual header instead of printing it to STDOUT (a library shouldn't print): read_table: Filter column 'nope' not found in the header of FILE; header is: 'PDB', 'score' 0.20 2026-07-05 CDT - addition of `ncol`, `nrow`, and `pnorm` functions - `filter` can filter by row names with `$_[1]` - `view` now accepts array of arrays in addition to AoH, HoA, and HoH [csort] - Two behavioural changes, both contained to the `csort` XSUB (the `cs_*` helpers are untouched). - Row names survive a Hash-of-Hashes sort. Sorting a HoH previously discarded the outer keys. Now each row is folded into a *fresh* row hash (a private container over aliased, read-only cells) that carries its outer key under a `row.name` column, so the name flows into whichever shape you request: my $hoh = { alpha => { id => 1 }, beta => { id => 2 } }; csort($hoh, 'id'); # AoH: each row gains a row.name field csort($hoh, 'id', 'hoa'); # HoA: an aligned row.name column - The column name defaults to `row.name` and can be overridden with an optional 4th argument (mirroring `hoa2hoh`'s named-key style): `csort($df, 'id', 'aoh', 'sample')`. - The outer key is authoritative — it wins over any pre-existing same-named field in the row. - Once present, the column is sortable like any other: `csort($hoh, 'row.name')`. - Because rows are now *copied* rather than shared, the caller's HoH is never mutated by the injection. (Minor behaviour change: output rows are no longer the same refs as the source rows.) - Clearer usage message. The signature is now `csort(...)`, so xsubpp no longer emits the misleading auto-generated `Usage: Stats::LikeR::csort(data, by, output=&PL_sv_undef)`. Argument count is checked by hand, and the croak now shows both real calling forms: Usage: csort($df, 'column.name', 'HoA') or csort($df, sub { $b->{'No.'} <=> $a->{'No.'} }, 'hoa') (optional 4th arg names the row-name column when sorting a HoH; default 'row.name') - `data`/`by`/`output` are read as `ST(0..2)`; `output` still defaults to matching the input shape. - Tightened validation messages. The `$data` croak now reads `hash-ref (HoA or HoH)`, and the `$by` croak includes a concrete example: `a column name (e.g. 'No.') or a comparator code-ref using $a and $b, e.g. sub { $b->{'No.'} <=> $a->{'No.'} }`. Existing HoA croaks (`unequal lengths`, `not found`, `not an array-ref`) are unchanged. - When sorting, undefined values in the sorting column are placed at the bottom [cor] - Fixed an unsigned-integer underflow in `kendall_tau_b` and added a regression test. - Bug: - In `kendall_tau_b`, concordant/discordant counts `C` and `D` are declared `size_t` (unsigned). The numerator was computed as: return (NV)(C - D) / denom; - The subtraction `C - D` happens in unsigned arithmetic *before* the cast to `NV`. When discordant pairs dominate (`D > C`), the result wraps to a huge positive value instead of going negative. - For the arrays: dG_kcal_mol: -7.765, -9.328, -10.326, -9.038, -9.608, -9.779, -9.975, -6.906 anomaly_rank: 154, 155, 161, 188, 76, 172, 173, 69 - there are `C = 9` concordant and `D = 19` discordant pairs (no ties). `9 - 19` wraps to `18446744073709551607`, so the function returned ~`6.6e17` instead of the correct `-10/28 = -0.3571428571`. - Fix: - Cast each operand to `NV` before subtracting, so the arithmetic is signed: return ((NV)C - (NV)D) / denom; - Only that one line changed. The denominator sums (`C + D + tie_x`, `C + D + tie_y`) are non-negative, so they were left as-is. - Regression test — `cor.t`: - Kendall on the offending arrays pinned to `-0.3571428571`. - Explicit `[-1, 1]` range guard (the real backstop — the pre-fix value `~6.6e17` blows past the bound regardless of exact magnitude), plus a negative-sign assertion. - Pearson (`-0.4889102301`), Spearman (`-0.4761904762`), and default-method coverage of the three `compute_cor` branches. - Kendall boundary cases: perfectly concordant (`+1`), perfectly discordant (`-1`), self-correlation (`+1`), and a tie case exercising `tie_x` in the denominator. - `no_leaks_ok` per method (guarded with `unless $INC{'Devel/Cover.pm'}`). - Croak paths: length mismatch, unknown method, zero-variance input. [XS refactor] - Consolidate helper functions to reduce binary size, find bugs, and back the changes with tests. Every change was validated by translating the XS (`ExtUtils::ParseXS`) and compiling the result with the module's own `ccflags`. - Outcome: - **Net change to the source:** ~154 fewer lines; helper-function count down by 4 (7 removed, 3 added). - **Genuine bugs fixed:** two instances of the same latent defect (see below). The rest of the work was behavior-preserving consolidation. - Function consolidation: - | Change | Before | After | |---|---|---| | Three-way `NV` comparator | `compare_rank`, `cmp_rank_item`, `cmp_rank_info`, `compare_NVs` | single `cmp_nv3` (reads the leading `NV` member, valid for `RankInfo`/`RankItem`/raw `NV`) | | Average-rank routine | `compute_ranks` + `compare_index` restoration sort | existing `rank_data` (scatters ranks into `out[idx]`, no second sort) | | String comparator | `cmp_string_wt`, `lm_str_qsort` (byte-identical) | single `cmp_string_wt` | | Multiplicity filter & set difference | `intersection` + `get_unique` (~90% shared); `Lonly`/`Ronly` duplicated bodies; a separate `set_difference()` | one shared `set_multiplicity()` with an "all vs. one" mode flag and a `from_last` flag: `intersection` (all), `Lonly` (one, first array), `Ronly` (one, last array) | - All merges were confirmed behavior-preserving: the collapsed comparators are equivalent on ordinary values, `NaN`, and infinities, and `compute_ranks` and `rank_data` produce identical average ranks. - Bugs: - Two comparators stabilized their sort by returning `a->idx - b->idx` directly, where the index field is an unsigned `size_t`. The subtraction wraps and is then truncated to `int`, which is implementation-defined and gives the wrong sign once a difference exceeds `INT_MAX`. - `compare_index` — removed entirely (the routine that used it, `compute_ranks`, was replaced by `rank_data`). - `cmp_pval` — the tie-break comparator in the p-adjust path. **Missed in the initial review; found later** via a `-Wconversion` compile of the earlier source. Fixed to compare with the `(a > b) - (a < b)` idiom. - Caveat on severity: on every mainstream ABI (LP64, LLP64, ILP32), the low-word truncation happens to reproduce the correct sign for any array smaller than ~2^31 elements, so this never produces a wrong result at realistic sizes. It is a portability/UB issue, not a runtime failure, which is why no functional test detects it (see "Testing", below). - `LikeR.xs` — consolidated helpers; `compare_index` removed; `cmp_pval` fixed. [`view`] - non-ASCII characters now print [`write_table`] - new option to output to LaTeX table 0.19 2026-07-01 CDT - numerous `SSize_t var1 = av_len(var) + 1` are changed to `size_t var1 = av_len(var) + 1` as `size_t`; as the result cannot be negative, in order to expand numerical range - Addition of `hoa2hoh`, `binom_test`, `chunk`, `get_union`, `get_unique`, `Lonly`, `Ronly`, `qcut`, and 3 tukey functions - Better warnings when non-array references are given to `intersection` - `view` now breaks columns into chunks for very wide data sets, more closely matching R's behavior 0.18 2026-06-28 CDT - `restrict` keyword added to numerous places within `intersection` to decrease CPU time - fix to dist.ini for dependencies - fixed POD rendering 0.17 2026-06-23 CDT (approx) - addition of `assign`, which adds new columns based on calculations from other columns - addition of `hoa2aoh`, transforming hash of arrays to array of hashes - addition of `predict`, using results from `aov`, `glm`, and `lm` - addition of `aoh2hoh` transforming array of hash into hash of hashes, `intersection`, `uniq`, and `vals` [`aov`] - Bug fixes: - **`size_t` underflow on empty arrays.** Three loops were bounded by `av_len(...)` compared against an unsigned counter; `av_len` returns `-1` for an empty array, which turned `k <= len` into a `SIZE_MAX` loop. The `stack()` value loop, the `.` column-expansion loop, and the `group.stats` column loop now use a signed `SSize_t` bound. - **HoH row count.** Row count for hash-of-hashes input was taken from the return value of `hv_iterinit`; it now uses `HvUSEDKEYS(hv)` with a separate `hv_iterinit`, matching `predict`. - **Buffer overflow in interaction parsing.** `strcpy(right, colon + 1)` into a fixed `char right[256]` is now `snprintf(right, sizeof(right), ...)`. - Performance / memory: - **Removed the per-row `row_x` scratch allocation.** Design rows are built directly into `X_mat[valid_n]`; `valid_n` simply does not advance on a rejected row. Interaction columns read their operands from the same in-progress row, so the logic is unchanged. - **`row_names` is no longer dead.** Surviving row names are transferred (pointer move, no copy) into `surv_names` to key `fitted.values`; rejected rows are freed in place. - **Dropped a `restrict` UB.** `orig_data_sv` aliases `data_sv`; the `restrict` qualifier was removed. - New, `predict`-compatible output keys: - **`coefficients`** — OLS estimates recovered by back-substitution on the R factor left in `X_mat` against Q'y in `Y` (no re-derivation). Keys are the expanded term names (`Intercept`, continuous names, `base.level` dummies, and `a:b` interaction products). Aliased columns are reported as `NaN`, which `predict` drops. - **`fitted.values`** — `Xb` over the non-aliased columns, keyed by surviving row name. Computed from a snapshot of the design (`Dsav`) taken before the QR overwrites `X_mat`. Costs one transient copy of the design matrix; negligible for typical ANOVA where the column count is small. - **`xlevels`** — sorted level list per factor, index 0 = reference, aligned with the contrast coding used to build the dummies. - **`family`** — `"gaussian"`. - Cleanup-path correctness: - `xlevels_hv`, `Dsav`, and `surv_names` are freed on both the "0 degrees of freedom" croak and the normal exit. The interaction-main-effects croak in PHASE 3 also frees `xlevels_hv`. - Known limitations (unchanged): - The intercept-stripping string surgery (`-1`, `+0`, `+1`, ...) operates on the whole RHS and can still mangle `I(x-1)`-style transforms; treat `I()` with arithmetic constants carefully. - Top-level keys `coefficients` / `fitted.values` / `xlevels` / `family` / `group.stats` share the return hash with the ANOVA rows; a predictor literally named one of those would collide. [`predict`] - New: factor-bearing interaction terms: - Previously, interaction coefficients such as `GroupB:Sexmale` or `GroupB:x` fell through to the continuous `evaluate_term` path and died on a nonexistent column. They are now handled directly: - **`dummy_hv`** stores each dummy's factor base index (an `IV`) instead of `&PL_sv_yes`, so a dummy name maps back to its `(base, level)` in O(1) (`level == name + strlen(base)`). `hv_exists` lookups are unaffected. - During coefficient caching, any `:` term with at least one factor-dummy component is routed to a separate list (`icopy` / `ibeta`); pure-continuous interactions (e.g. `x:z`) stay on the existing `evaluate_term` path, so prior behavior is preserved. - Each routed term is parsed once into flat component arrays. Factor components store a base index and level pointer; continuous components store the term string and get the same up-front column-existence validation as main terms. - Per row, each factor's raw level is read once into `raw_lv[]` and reused by both main effects and interactions (no duplicate `get_data_string_alloc`). An interaction's value is the product of its components: a factor component contributes `1.0` iff the row's level matches the dummy's level (reference levels give `0`), continuous components go through `evaluate_term`. - This covers factor×factor, factor×continuous, continuous×continuous, and n-way combinations. - Other: - HoH row count uses `HvUSEDKEYS` (already present). - The unseen-factor-level croak now frees every level string already read for the current row, not just the current one. [Tests] - **`aov.t`** — one-way ANOVA against hand-computed values (Df / Sum Sq / Mean Sq / F / decomposition); identical results across HoA / HoH / AoH / stacked input; simple regression; `.` expansion; intercept removal (`-1`); two-way with interaction (Type I SS on a balanced design); NaN listwise deletion; all croak paths; leak checks. - **`predict.t`** — `predict(training) == fitted.values` round-trips for one-way, regression, factor×factor, factor×continuous, and continuous×continuous models; explicit predicted values; agreement across HoA / AoH / HoH / flat newdata; no-newdata path; binomial `link` vs `response`; gaussian identity link; all croak paths; leak checks. - Leak tests use `no_leaks_ok` guarded by `unless $INC{'Devel/Cover.pm'}` and skipped when `Test::LeakTrace` is absent. - Assumptions worth confirming: - The NaN-deletion test relies on `evaluate_term` returning `NaN` for a non-finite response value (an `Inf - Inf` NaN is fed in deterministically). - The continuous×continuous round-trip relies on `evaluate_term("x:z")` yielding `x * z` — the same assumption the pre-existing `predict` continuous-interaction path already made. If that path was untested, this round-trip now exercises it. [`view`] - now returns colored output; fixed bug with incorrect widths; undefined values show as `undef` rather than `NA`, as in Data::Printer [`csort`] - now accepts Hash of Hashes; addition of `restrict` which should decrease calculation time [filter] - **Added hash-of-hashes (HoH) input.** In addition to AoH and HoA, `filter` now accepts an HoH (`{ key => { col => val, ... }, ... }`); each inner hash is one row, and matching keys are preserved by default (HoH -> HoH). - **Added `output.type`.** `filter($df, $pred, 'output.type' => 'aoh'|'hoa')` selects the returned shape (aliases `out` / `output_type`; a bare positional type also works). When omitted, the input shape is preserved. `hoh` is not a selectable output, since it would require choosing a key column. - **`col()` reworked, not removed.** Both predicate forms are kept: `col('age') >= 18` still works and is the concise/composable option, while a coderef covers everything else. Internally `col()` is now **pure Perl** — an overloaded class that builds a per-row closure — and `filter` unwraps that closure so `col()` and a coderef share one evaluation path. The previous standalone XS predicate evaluator (`filt_eval`/`filt_ctx`) is gone; delete it if your tree still has it. One consequence: a `col()` comparison now costs the same per row as the equivalent coderef (a Perl call), rather than being evaluated in C. - **Unchanged guarantees:** the input frame is never modified; `undef` (and, for numeric ops, non-numeric) cells never match a `col()` comparison; AoH/HoH rows are shared rather than copied where possible; keep-all/keep-none shapes are well defined per output type; Perl 5.10 compatibility is retained. A latent `SvTRUE(POPs)` double-evaluation in the per-row call helper (which crashed on perls where `SvTRUE` is a multi-eval macro) was fixed along the way. [read_table] - Added an opt-in `auto.row.names` argument so `read_table` can read the file R produces by default from `write.table(x, sep="\t")`. - The problem: - R's `write.table` defaults to `row.names=TRUE, col.names=TRUE`, which writes the row-names column in every data row but emits no header label for it. So a frame with N columns comes out as N header fields over N+1 data fields — e.g. `mtcars` gives 11 headers but 12-field rows. By default `read_table` (correctly) rejects that as ragged: Alignment error on mtcars.tsv data row 1 (12 fields vs 11 headers). - The change: - `auto.row.names` turns on R's own `read.table` rule: **when, and only when, the header is exactly one field short of the data rows, treat the first field of each row as an (unlabelled) row-names column.** # default: the leading column is named 'row_name' my $df = read_table('mtcars.tsv', 'auto.row.names' => 1); # or give it a name my $df = read_table('mtcars.tsv', 'auto.row.names' => 'model'); - The synthesized column behaves like any other first column: it appears in `aoh` and `hoa` output, and for `hoh` it becomes the default key (so rows are keyed by the model name). This also lines up with the existing handling of R's `col.names=NA` output (a blank leading header), which still produces a `row_name` column with no flag needed. - What did not change: - The strict alignment check is still the default. Without `auto.row.names` the lopsided file still croaks, and even with it, a row that is off by anything other than exactly one field still croaks — so the corruption guard only relaxes for the one case R itself treats specially. - Tested in `t/read_table.2.t` (16 assertions, Perl 5.10.1 and 5.38): aoh / hoa / hoh output, custom column name, the already-aligned file (flag is a no-op), the `col.names=NA` path, and the strict / ragged croak paths. - additional bugfix: # This is a comment id,name,val 1,Alice,10.5 2,Bob, 3,Charlie,15.2 - would not be read correctly using `read_table`, but now is read correctly [value_counts] - now accepts array of hashes 0.16 2026-06-17 CDT - changes to dist.ini, the minimum Perl version disappeared when I fixed other problems - clarifications between run time and test dependencies - addition of `csort` function to sort AoH and HoA - addition of `aoh2hoa` to translate array of hashes into a hash of arrays - fix of long double functions: https://www.cpantesters.org/cpan/report/5d5d9836-6a5f-11f1-aadb-63fd6d8775ea [`glm`] - output residual keys now use names, not integers [`lm`] [Bug fixes] - Memory leak on the zero-degrees-of-freedom error path. When `valid_n <= p`, the cleanup freed the `valid_row_names` *array* but not the per-row name strings it held (those had been transferred out of `row_names`, whose own array was already freed). The strings leaked on every such error. Added the per-entry `Safefree` loop before freeing the array, matching the normal path. - HoH input validated only the first row. Only the first hash value was checked to be a `HASHREF`; subsequent values were `SvRV`'d unconditionally, so a malformed row (`{ a => {...}, b => 5 }`) dereferenced a non-reference. Every row is now validated, with the partial allocations cleaned up before the `croak`, mirroring the existing AoH path. - `isspace` on a possibly-signed `char`. `isspace(*src)` is undefined for byte values ≥ 0x80 on platforms where `char` is signed. Cast to `(unsigned char)` before the call. [Speed / RAM improvements] - Formula buffer is now heap-allocated to fit. `char f_cpy[512]` silently truncated any longer formula. Replaced with a buffer sized to `strlen(formula) + 1`, so there is no fixed limit and no truncation. - `.`-expansion buffer is now a growable heap buffer. `char rhs_expanded[2048]` silently dropped expanded terms once full. It is now a buffer that doubles on demand. Appends also went from `strcat` (which rescans from the start every time — O(n²) over many columns) to an O(1) amortised append that tracks the write position. - No more per-row scratch allocation in matrix construction. The original `safemalloc`'d a `row_x` buffer, filled it, copied it into `X`, and freed it *for every row* — `n` allocations plus `n*p` copies. Each candidate row is now written straight into `X` at its prospective commit slot; a row that fails listwise deletion is simply overwritten by the next candidate. This removes the `n` allocate/free cycles and the copy loop entirely. - Categorical levels sorted with `qsort`. The level list used an O(n²) bubble sort; replaced with `qsort` (relevant only for high-cardinality factors). - Unused tail of `X` reclaimed after listwise deletion. `X` is allocated for all `n` rows up front (`valid_n` is unknown until rows are scanned). When rows are dropped, `X` is now `Renew`ed down to `valid_n * p`, returning the unused tail to the allocator before the OLS phase. - Minor robustness. The argument-parsing index was widened from `unsigned short` to `I32` to match `items`, and the HoH row count now uses `HvUSEDKEYS` rather than relying on `hv_iterinit`'s return value. [Known limitations (left unchanged)] - A multi-way term such as `a*b*c` is split only on the first `*`, so it yields `a`, `b*c`, and `a:b*c` rather than a full three-way expansion. Deeper interactions silently fail (the unparsable term evaluates to `NaN` and the rows are dropped). This matches the documented two-way `*` support. - HoA input takes the row count from the first column; columns shorter than that simply contribute dropped rows rather than raising an error. [`oneway_test`] - Bug fixes: - Memory leaks on error paths. Nearly every `croak` after an allocation leaked memory. `croak` does a `longjmp`, so anything allocated but not yet freed is lost. Affected paths: - AoA and hash first-pass errors leaked `sizes` and any `gnames[]` entries allocated so far. - Formula-mode "not found as an array ref" errors leaked `lhs` and `rhs`. - All post-allocation errors now route through a single `fail:` label that frees every pointer unconditionally. Pointers are initialised to `NULL` and `gnames` is zero-allocated with `Newxz`, so the cleanup is always safe to run. - Undefined and non-numeric cells silently coerced to `0.0`. The original second pass used `(svp && *svp) ? SvNV(*svp) : 0.0`, meaning an `undef` or non-numeric cell was quietly treated as zero, silently corrupting the F-statistic. Each cell is now validated with `SvOK` and `looks_like_number`; the call dies naming the group and observation index, consistent with the rest of `Stats::LikeR` (`mean`, `sum`, `cor`, etc.). - Unsigned wraparound on empty array input. `k = (size_t)av_len(in_av) + 1` cast to `size_t` *before* adding, so an empty array (`av_len` returns `-1`) produced `SIZE_MAX` rather than `0`. Changed to `k = (size_t)(av_len(in_av) + 1)` so the `+1` is done in signed arithmetic before the cast. - Unreliable group count from `hv_iterinit`. `hv_iterinit` returns the number of buckets in use rather than the number of keys for tied hashes. Replaced with `HvUSEDKEYS`, which always returns the correct key count. - Improvements: - `var.equal` accepted as an alias for `var_equal`. R users write `var.equal`; the argument parser now accepts both spellings. - Perl memory API used throughout. `safemalloc` and manual `memcpy` replaced with `Newx`, `Newxz`, `savepv`, and `savepvn`. `savepvn` additionally preserves embedded NUL bytes in group key strings, which the previous `strlen`-based copies silently truncated. - Known limitations (not changed): - A factor column named `Residuals` or `group.stats` in a formula call will collide with reserved top-level keys in the result hash. - Group names containing an embedded NUL are stored correctly but are still truncated at `strlen` when written into the output hash keys. [`view`] - default view shifted to 80 characters to match Linux window length - New features: - **`rows` is accepted as a synonym for `n`** (the number of rows shown). Passing both `n` and `rows` is an error. - **Unknown arguments are now rejected.** `view` validates its argument names against the documented set (`n`, `rows`, `na`, `max_width`, `ellipsis`, `gap`, `cols`, `columns`, `to`, `return_only`, `row.names`, `row_names`) and dies listing any it does not recognise, so a misspelt option (e.g. `widht`) is caught instead of silently ignored. - **`n` / `rows` is validated.** It must be a non-negative integer; `undef` or a non-numeric value now dies with a clear message instead of producing warnings and being treated as `0`. - **flat/simple hashes are accepted as input** - Bug fixes: - **`n => 0` now still prints the column header.** Column names were collected only from the rows being shown, so requesting zero rows produced an empty header line. At least one row is now scanned (when data exists) so the header always lists the columns. - **An empty hash (`{}`) no longer dies.** It was rejected as *"neither ARRAY nor HASH"*; it is now shown as an empty table (`0 rows x 0 cols`), matching the handling of an empty array. - **The `row_names` alias now drives the Hash-of-Hashes label header.** The header for the row-label column consulted only `row.names`, so `row_names => 'id'` displayed `row_name` instead of `id`. Both spellings are now honoured consistently. - **Malformed nested values degrade gracefully.** A Hash-of-Arrays column or Hash-of-Hashes row whose value is not actually an array/hash reference now renders as empty cells rather than throwing a dereference error. - Performance: - Column gathering no longer sorts once per scanned row. Unique column names are collected across the scanned rows and sorted a single time (same output order), and the ellipsis length is computed once rather than per cell. - Tests: - `t/view.t` is self-contained (the `view` implementation is inlined; it loads no other files) and covers the new argument handling, the bug fixes above, and the existing AoH / HoA / HoH behaviour, alignment, truncation, and output-path handling. [`wilcox_test`] - Corrected four bugs in the `wilcox_test` XSUB plus a portability fix in its exact signed-rank helper. Behaviour on valid input is unchanged: the R-agreement cases (unpaired `W = 58`, `p = 0.13292`; paired one-sided `V = 40`, `p = 0.019531`; separated exact `W = 0`, `p = 0.028571`) all still match R's `wilcox.test`. - Bug fixes: - **Invalid `alternative` is now rejected.** Any value other than `less` or `greater` previously fell through to the two-sided branch and returned a two-sided result mislabelled with the bad string, so a typo like `alternative => "twosided"` silently "worked". It now croaks unless `alternative` is one of `two.sided`, `less`, `greater`. - **Zero/negative variance is guarded.** When every observation is tied the approximation's variance collapses to 0 and the old code divided by `sqrt(0)`: `wilcox_test([5,5,5], [5,5,5])` returned `p = 0` (a "significant" difference between identical samples). It now warns and returns `p = 1`. - **Two-sided continuity correction at `z = 0`.** R uses `sign(z) * 0.5`, so the correction is `0` when the statistic sits exactly on its mean; the old code used `-0.5`. Example: `wilcox_test([1,4], [2,3], exact => 0)` changed from `p = 0.698535` to `p = 1` (matches R). - **`exp` no longer shadows libm.** The local `exp` accumulator (mean of the statistic) shadowed the C library `exp()`; renamed to `mean_w` (two-sample) and `mean_v` (signed-rank). No active miscompute, removed as a latent hazard. - Cosmetic: - Collapsed a no-op ternary that assigned the same signed-rank exact method string on both branches; the `method` field is now simply `Wilcoxon signed rank exact test`. - Portability (exact signed-rank helper): - **`exact_psignrank` no longer calls `powl()`.** The `2^n` normaliser is now built by exact repeated doubling, which has no long-double libm dependency. This fixes an `Undefined symbol "powl"` load failure reported by a CPAN smoker (FreeBSD, perl 5.20, `nvtype=double`) whose libm lacks the long-double math functions; the symbol resolved on glibc, which is why local builds passed. `long double` accumulation in the DP is retained — only the `powl` call was at fault. - **`int` → `size_t`** for `n`, `max_v`, and the DP loop counters, which also removes a `size_t`-to-`int` narrowing at the call site. The `floor()` result (`k`) stays signed so its negative-`q` sentinel still fires, and is cast to `size_t` only after the `k < 0` check. - Tests: - Added `t/wilcox_test.t` (flat, no subtests): R-agreement cases, option handling (`paired`, `correct`, `exact`, `mu`, named/positional `x`/`y`, NA dropping), regressions for all four bug fixes, argument-error and `alternative`-validation checks, output shape, and `no_leaks_ok` coverage of the two-sample, exact, and paired allocation paths. 0.15 2026-06-11 CDT - `view` function added, similar to R's `head` - `read_table`: filter => { 'Testosterone, total (nmol/L)' => sub { defined $_ }, } - was broken by the change in undefined variables in 0.14, but is back to being `undef` - `col2col` improvement in sectioning in README - Numerous changes to prevent quadmath/long double CPAN test failures - Minimum Scalar::Util version in dist.ini is now 1.22, see https://www.cpantesters.org/cpan/report/6b682236-6567-11f1-a3bc-a055f9c4ba34 - `Digest::SHA` removed as a dependency [`read_table`] - Bug fixes: - **A comment-prefixed header is now read correctly.** `read_table` strips a leading comment marker from the header line (so a file may begin with `#id,val`), but that strip was dead code: the XS parser skipped *every* line beginning with the comment string before the callback ever saw it, so a commented header was silently dropped and the first data row was mistaken for the header. The parser now delivers the first content line even when it begins with the comment marker, and only skips comment lines after the header has been seen. - **Carriage returns inside quoted fields are preserved.** The parser stripped `\r` unconditionally, so a quoted value such as `"x\ry"` lost its carriage return and would not survive a `write_table` -> `read_table` round-trip. `\r` is now stripped only as part of a trailing CRLF line ending and as a stray CR *outside* quotes; inside quotes it is literal data. - **Duplicate column names no longer corrupt `hoa` output.** With `output.type => 'hoa'`, a repeated column name pushed the same cell once per occurrence, so the affected columns came out longer than the others and the arrays no longer lined up by row. Columns are now keyed by unique header name (first-seen order preserved, later values win, one warning emitted). - **A defined non-CODE callback is now an error.** Passing a defined argument that was not a CODE reference silently fell through to slurp mode and ignored the argument; it now croaks (*"callback must be a CODE reference"*). - **An undefined/empty `hoh` row-name now dies instead of keying on `""`.** With `output.type => 'hoh'`, a row whose row-name column was empty/undef was stored under the `''` key and raised *"uninitialized value"* warnings. It now dies, naming the column and the offending data row. - **A numeric filter key past the last column now dies.** A 1-based numeric filter key greater than the column count was accepted, then silently extended every row through the `$_` write-back. It is now rejected up front with a message naming the column count. - **`sep` and `delim` together now die.** Supplying both silently preferred `delim`; passing both is now an explicit error (`delim` remains an alias for `sep` when used alone). - **The library no longer prints to STDOUT.** The unknown-argument path used `say` to dump the offending names to STDOUT before dying; the names are now carried in the `die` message itself. - Better diagnostics: - Alignment errors now report **which data row** is ragged (*"Alignment error on FILE data row N (X fields vs Y headers)"*), instead of only the field/header counts. - Memory-leak fixes (exception paths): - The parser allocated its working buffers (`current_row`, `field`, and — in slurp mode — `data`) in the XS `INIT:` block, i.e. *before* any validation, and freed them only by falling off the end of the function. Any non-local exit therefore leaked: - the open-failure `croak` leaked the row buffer and field (and the slurp accumulator); - far more commonly, a `die` thrown **inside the row callback** — which `read_table` does routinely on alignment errors, bad row names, and filter exceptions — unwound straight out of the XS frame and leaked the field, the current row, the line buffer, the slurp accumulator, *and the open file handle*. - Allocations now happen in `CODE:` after every croak-able check, and every long-lived resource (the file handle via `SAVEDESTRUCTOR_X`, the buffers via `SAVEFREESV`) is tied to the save stack, which an exception unwinds. Measured with `Test::LeakTrace`: a `die` mid-file went from 5 leaked SVs to 0, and an open failure from 2 to 0. This is the likely source of the constant-size leaks seen in CPAN-tester reports for the exception-path tests. - Performance: - **~2.5x faster parsing** (57 -> 145 MB/s on a 100k-row quoted file). The core loop appended one character at a time with `sv_catpvn(field, &ch, 1)`; it now scans runs of ordinary bytes with `memchr` / a bounded scan and appends each run in a single `sv_catpvn`, copying field contents in bulk rather than byte by byte. - Internal / non-behavioral: - XS declarations moved from `INIT:` to `PREINIT:`; allocations deferred into `CODE:` (see the leak fixes above). - The filter loop now aliases the row hash with `local *_ = \%line_hash` instead of copying it with `local %_ = %line_hash`. This removes a full per-row hash copy for every filtered row and fixes a latent staleness bug: after a filter mutated `$_` and the change was written back, `%_` still reflected the pre-mutation copy, so a subsequent filter in the same row saw stale values. With aliasing, `%_` *is* the row, so write-backs are always visible. - Known limitation (not changed): - **`undef.val` does not round-trip back to `undef`.** `write_table` renders an `undef` cell as an empty field by default, and `read_table` maps an empty field back to `undef`, so the *default* round-trip is clean. But if a file is written with a token such as `'undef.val' => 'NA'`, `read_table` has no inverse option and reads `NA` back as the string `'NA'`. `read_table` also cannot distinguish a deliberately quoted empty string (`""`) from a missing value -- both become `undef`. Adding an `na.strings`-style option to `read_table` (mapping configurable tokens and/or empty fields to `undef`) would close this gap. [`write_table`] - Behavior change: - **`undef` cells now write as an empty field, not an empty string.** A missing or `undef` value renders as nothing between separators (`a,,c`) rather than a quoted empty string (`a,'',c` / `a,"",c`). Supplying `'undef.val' => 'NA'` (or any other token) still overrides this, exactly as before. This is the only change that can alter the bytes of an existing output file; if you relied on the previous default, pass `'undef.val' => ''` to keep an explicit empty field, or your chosen placeholder. - Bug fixes: - **Wide-character / UTF-8 column names and row keys now round-trip.** Previously, cells were looked up with the raw bytes of the column name (`hv_fetch(..., SvPV_nolen(name), strlen(name), ...)`), which fails to match a UTF-8-flagged hash key: the column header printed correctly but every cell under it came back empty. All lookups now fetch by SV (`hv_fetch_ent`), header lists are gathered and sorted as SVs (`sortsv` + `sv_cmp`, preserving the flag) instead of being round-tripped through `char *`, and the `row.names` column is matched with `sv_eq` rather than `strcmp`. Embedded NUL bytes in keys are handled correctly as a side effect. - **`col.names => []` no longer loops forever.** An empty `col.names` array made `av_len()` return `-1`, which — compared against an unsigned `size_t` loop index — wrapped to `SIZE_MAX` and ran effectively without end. This was fixed for flat hashes previously; it was still present for hash-of-hashes, hash-of-arrays, and array-of-hashes, plus both `row.names` header-filtering loops. All such loops now use a signed index. - **Tables wider than 65,535 columns no longer hang.** One header loop used an `unsigned short` index that silently wrapped past 65,535 and never terminated. It now uses `size_t` like the rest of the code. - **Flat-hash cells holding a reference now croak.** Every other input shape rejects a nested reference with *"Cannot write nested reference types to table"*; a flat hash instead stringified it (e.g. `ARRAY(0x55...)`) into the file. It now croaks consistently. - **`'undef.val' => undef` is handled cleanly.** It previously called `SvPV_nolen` on `undef`, raising an *"uninitialized value"* warning and yielding an empty string by accident. It is now treated explicitly as an empty field, with no warning. - Memory-leak fixes (exception paths): - The row-key list gathered for hash-of-hashes input was leaked when the output file could not be opened. - The *"Could not get headers"* croak on hash-of-arrays input leaked both the already-open filehandle and the headers array. - Internal / non-behavioral: - Numeric row labels are now formatted into a reused stack buffer instead of a per-row `savepv()` / `safefree()` allocation (no functional change; removes a cast-away-`const` and one allocation per row). - Several signed/unsigned index types were made consistent (`SSize_t` vs `size_t`) to match `av_len()` and silence the conditions behind the loop bugs above. - Tests: - `t/write_table.t` expanded from 17 to 69 assertions. New coverage targets each fix above: the empty-field default and `undef.val => undef` (no warning), `col.names => []` termination across all four input shapes, the >65,535-column header loop (gated behind `EXTENDED_TESTING=1`), in-sequence numeric row labels, nested-reference rejection, CSV quoting corners (carriage return, separators inside column names, multi-character separators), empty input writing no file, and UTF-8 column names and row keys. Two leak assertions cover the exception paths above. 0.14 2026-06-08 CDT - `filter` function added for rows - `read_table` reads undefined values to `undef` instead of `NA`, which makes calculations easier - `write_table` writes undef by default as an empty string `''` - `hoh2hoa` transforms a hash of hashes into an hash of arrays - `quantile` uses `NV` instead of `double` to allow for high-precision 128-bit floats to be used on quadmath machines when available: https://www.cpantesters.org/cpan/report/296f4868-631f-11f1-abba-ff15558d240b - Numerous switches from `double` to `NV` for local precision, like above - numerous changes to `col2col` for ease of use and working with datasets with numerous undefined values - dist.ini now links to math library when compiling: https://www.cpantesters.org/cpan/report/785e26d8-6397-11f1-89c0-dc066e8775ea - `fisher_test` now should be complete, errors with confidence intervals fixed 0.13 2026-06-07 CDT - `read_table`: speed improvements; commented headers are now allowed - `write_table`: fix for Attempt to free temp prematurely: SV 0x56417a2ae610 at t/write_table.t line 182. main::wrote_ok(",age\x{a}Alice,30\x{a}Bob,25\x{a}", "row.names => 'name' uses that column as labels", HASH(0x56417a272250), "row.names", "name") called at t/write_table.t line 203 Attempt to free unreferenced scalar: SV 0x56417a2ae610 at t/write_table.t line 183. main::wrote_ok(",age\x{a}Alice,30\x{a}Bob,25\x{a}", "row.names => 'name' uses that column as labels", HASH(0x56417a272250), "row.names", "name") called at t/write_table.t line 203 - `write_table` gives better warnings for incorrect types of data given - Numerous changes to dist.ini to improve CPAN testing, especially for Win32 0.12 2026-06-08 CDT - `add_data` can also take hash of arrays, and various mixes of data types - `ljoin`: Addition of `restrict` keywords in many places; should improve CPU performance - Better POD formatting, correction of output hash for README's `add_data` - `chisq_test` can now accept hash of hashes as input - new `transpose` function for switching 2D hash keys and 2D array indices, and `col2col` for comparing columns against columns - removed unused function from C helpers - `value_counts`: addition of restrict keywords in preinit, should improve CPU performance - MANIFEST.skip changed to MANIFEST.SKIP to improve CPAN testing - using `is_deeply` for tests of `transpose`, which may or may not work with CPAN testers (experimental) - Added function name to warnings, so I actually know which function is producing the error - `write_table` can also take `file` and `data` as args, in addition to positions - fixed `write_table` as it could hang if given empty `col.names` or `row.names` - Added `__EXTENSIONS__` to source XS file for better CPAN testing 0.11 2026-06-03 CDT - better POD formatting for tables - addition of MANIFEST.skip to get better testing results on CPAN - `glm`: bugfix for when there is no intercept in the formula, new test cases in t/glm.t - `write_table` now accepts simple hashes as input, in addition to hash of arrays, hash of hashes, and arrays of hashes - Better documentation for t-test 0.10 2026-06-01 CDT (approx) - changes to compilation for CPAN, trying to get this work on Windows - Addition of `prcomp` and `value_counts` - `matrix` will work without key names, just like in R. Testing for `matrix` has improved. 0.09 2026-06-01 CDT (approx) - context changes in XS `dTHX`, `pTHX_`, and `aTHX_` to get better CPAN testing results - `restrict` keywords added to `lm` to increase speed 0.08 2026-05-26 CDT - Speed improvement in `summary` of hashes. - Addition of `add_data`, `dnorm`, `group_by`, `ljoin`, and `mode` functions - Chi-squared function no longer has Perl wrapper, and all code is in XS, which should result in a minor speed increase with 1 less function call. - Compiler changes for GNU source and inclusion of `strings.h`, to ensure more CPAN testing works better. - `read_table` now returns hash-of-hash in {row}{column} 0.07 2026-05-24 CDT - Addition of `summary` function. - Formulas can now be omitted from `aov`, resulting in a stacked calculation as R would think. - Addition of `oneway_test` for multi-group comparisons that does not assume normality like `aov` does. - `read_table` and `write_table` now automatically set separators for `.csv` files as `,` and `.tsv` files as `"\t"`, respectively, so these values no longer need to be specified separately from the file name. 0.06 2026-05-19 CDT - Changed compiler options so that Solaris will work - signed integers changed to unsigned in `glm` - Added restrict keywords to `power_t_test`, and made `int` to `unsigned int` 0.05 2026-05-08 CDT - Leak testing for `sample` - removal of Data::Printer dependency for easier CPAN testing - switched several `unsigned int` variable to `I32` so that clang doesn't complain - added restrict keyword for `sample` 0.04 2026-5-17 CDT - addition of `sample` function - GNU source, to maximize compatibility and ease installation - removal of JSON dependency to ease installation 0.03 2026-5-13 CDT - Compatibility back to Perl 5.10 0.02 2026-5-7 CDT - back-compatible to Perl 5.10, instead of original 5.40, ensuring more people can use it - added var_test - mean, min, sum, median, var, and max die with undefined values, and print the offending indices - "group.stats" added to aov, for TukeyHSD in the future - "cor" dies when given data with standard deviation of 0 - `write_table` now has `undef.val` option, which shows how undefined values are printed to tables, which is `NA` by default.