A reality-near demonstration article exercising the full JATS feature set

Astrid M. Lindqvist1,2, Chidi E. Okonkwo2, Claire Beaumont3
Abstract

Background

This article demonstrates a complete JATS feature set for rendering across layouts: structured affiliations with ROR identifiers, ORCID, a corresponding author, a numbered table of contents, a multi-row table header, a figure, a display equation, dense footnotes, and a structured reference list.

Methods

We rendered the article through the production layout engine and validated each feature against the resulting PDF and self-contained HTML.

Results

All showcased elements rendered correctly, including the affiliation ROR links and the cross-referenced citations and footnotes.

Keywords: JATS, scholarly publishing, PDF rendering, ROR

Introduction

Structured markup separates content from presentation, which is the precondition for reliable automated typesetting. [1] We observed that affiliation handling, in particular, varies widely between publishers and even between articles. Tables with spanning headers require the serializer to preserve colspan and rowspan faithfully.1

Each element carries explicit semantics, so a downstream renderer can make correct layout decisions without heuristics. [2] The reference list is where small inconsistencies accumulate into visible defects if the renderer is not defensive. Mathematical content must survive the round trip from MathML to the rendered output without namespace loss.2

In practice, the hardest cases are the ones where real-world documents deviate from the idealized schema. [3] Footnotes interact with pagination in ways that are easy to get wrong when content is dense. Cross-references should resolve to stable anchors so that both print and screen output remain navigable.3

We observed that affiliation handling, in particular, varies widely between publishers and even between articles. [4] A two-column layout stresses the flow algorithm far more than a single column, especially around floats. Structured markup separates content from presentation, which is the precondition for reliable automated typesetting.4

Background

The reference list is where small inconsistencies accumulate into visible defects if the renderer is not defensive. [5] Tables with spanning headers require the serializer to preserve colspan and rowspan faithfully. Each element carries explicit semantics, so a downstream renderer can make correct layout decisions without heuristics.5

Footnotes interact with pagination in ways that are easy to get wrong when content is dense. [6] Mathematical content must survive the round trip from MathML to the rendered output without namespace loss. In practice, the hardest cases are the ones where real-world documents deviate from the idealized schema.6

A two-column layout stresses the flow algorithm far more than a single column, especially around floats. [7] Cross-references should resolve to stable anchors so that both print and screen output remain navigable. We observed that affiliation handling, in particular, varies widely between publishers and even between articles.7

Tables with spanning headers require the serializer to preserve colspan and rowspan faithfully. [8] Structured markup separates content from presentation, which is the precondition for reliable automated typesetting. The reference list is where small inconsistencies accumulate into visible defects if the renderer is not defensive.8

Methods

Mathematical content must survive the round trip from MathML to the rendered output without namespace loss. [9] Each element carries explicit semantics, so a downstream renderer can make correct layout decisions without heuristics. Footnotes interact with pagination in ways that are easy to get wrong when content is dense.9

Cross-references should resolve to stable anchors so that both print and screen output remain navigable. [10] In practice, the hardest cases are the ones where real-world documents deviate from the idealized schema. A two-column layout stresses the flow algorithm far more than a single column, especially around floats.10

Structured markup separates content from presentation, which is the precondition for reliable automated typesetting. [11] We observed that affiliation handling, in particular, varies widely between publishers and even between articles. Tables with spanning headers require the serializer to preserve colspan and rowspan faithfully.11

Each element carries explicit semantics, so a downstream renderer can make correct layout decisions without heuristics. [12] The reference list is where small inconsistencies accumulate into visible defects if the renderer is not defensive. Mathematical content must survive the round trip from MathML to the rendered output without namespace loss.12

Table 1 Primary and secondary outcomes by treatment group. The header spans two rows.
GroupPrimary outcomeAdverse events
Recovery %95% CI
Placebo4134–483
Low dose5851–655
High dose7266–787
Control4942–562
d=μ1μ2σ12+σ222 ((1))
Figure 1 Recovery rate by treatment group. The high-dose group (green) showed the highest recovery.

Study design

In practice, the hardest cases are the ones where real-world documents deviate from the idealized schema. [13] Footnotes interact with pagination in ways that are easy to get wrong when content is dense. Cross-references should resolve to stable anchors so that both print and screen output remain navigable.13

We observed that affiliation handling, in particular, varies widely between publishers and even between articles. [14] A two-column layout stresses the flow algorithm far more than a single column, especially around floats. Structured markup separates content from presentation, which is the precondition for reliable automated typesetting.14

The reference list is where small inconsistencies accumulate into visible defects if the renderer is not defensive. [15] Tables with spanning headers require the serializer to preserve colspan and rowspan faithfully. Each element carries explicit semantics, so a downstream renderer can make correct layout decisions without heuristics.15

Footnotes interact with pagination in ways that are easy to get wrong when content is dense. [16] Mathematical content must survive the round trip from MathML to the rendered output without namespace loss. In practice, the hardest cases are the ones where real-world documents deviate from the idealized schema.16

Data collection

A two-column layout stresses the flow algorithm far more than a single column, especially around floats. [17] Cross-references should resolve to stable anchors so that both print and screen output remain navigable. We observed that affiliation handling, in particular, varies widely between publishers and even between articles.17

Tables with spanning headers require the serializer to preserve colspan and rowspan faithfully. [18] Structured markup separates content from presentation, which is the precondition for reliable automated typesetting. The reference list is where small inconsistencies accumulate into visible defects if the renderer is not defensive.18

Mathematical content must survive the round trip from MathML to the rendered output without namespace loss. [19] Each element carries explicit semantics, so a downstream renderer can make correct layout decisions without heuristics. Footnotes interact with pagination in ways that are easy to get wrong when content is dense.19

Cross-references should resolve to stable anchors so that both print and screen output remain navigable. [20] In practice, the hardest cases are the ones where real-world documents deviate from the idealized schema. A two-column layout stresses the flow algorithm far more than a single column, especially around floats.20

Statistical analysis

Structured markup separates content from presentation, which is the precondition for reliable automated typesetting. [21] We observed that affiliation handling, in particular, varies widely between publishers and even between articles. Tables with spanning headers require the serializer to preserve colspan and rowspan faithfully.21

Each element carries explicit semantics, so a downstream renderer can make correct layout decisions without heuristics. [22] The reference list is where small inconsistencies accumulate into visible defects if the renderer is not defensive. Mathematical content must survive the round trip from MathML to the rendered output without namespace loss.22

In practice, the hardest cases are the ones where real-world documents deviate from the idealized schema. [1] Footnotes interact with pagination in ways that are easy to get wrong when content is dense. Cross-references should resolve to stable anchors so that both print and screen output remain navigable.1

We observed that affiliation handling, in particular, varies widely between publishers and even between articles. [2] A two-column layout stresses the flow algorithm far more than a single column, especially around floats. Structured markup separates content from presentation, which is the precondition for reliable automated typesetting.2

Acknowledgments

We thank the demonstration reviewers for their feedback.

References

1.Maier B, Schmidt C. Structured publishing workflows for small presses. J Scholarly Publ. 2019;4:13-24. 10.1000/demo.2019.4.13
2.Schmidt C, Rossi D. Metadata interoperability in JATS. Inf Stand Q. 2020;5:16-27. 10.1000/demo.2020.5.16
3.Rossi D, Okonkwo E. Effect sizes in small randomized trials. Trials Today. 2021;6:19-30. 10.1000/demo.2021.6.19
4.Okonkwo E, Beaumont F. A survey of PDF rendering pipelines. Digit Libr Rev. 2022;7:22-33. 10.1000/demo.2022.7.22
5.Beaumont F, Lindqvist G. Affiliation disambiguation with ROR. Scholarly Kitchen. 2023;8:25-36. 10.1000/demo.2023.8.25
6.Lindqvist G, Hassan H. Hanging indents considered carefully. J Doc Eng. 2024;9:28-39. 10.1000/demo.2024.9.28
7.Hassan H, Nakamura I. Two-column flow in paged media. J Scholarly Publ. 2025;10:31-42. 10.1000/demo.2025.10.31
8.Nakamura I, Petrov J. Footnote placement strategies. Inf Stand Q. 2018;11:34-45. 10.1000/demo.2018.11.34
9.Petrov J, Ferreira K. MathML preservation across converters. Trials Today. 2019;12:37-48. 10.1000/demo.2019.12.37
10.Ferreira K, Kowalski L. Anchor stability in long documents. Digit Libr Rev. 2020;13:40-51. 10.1000/demo.2020.13.40
11.Kowalski L, Andersen M. Reference label styles compared. Scholarly Kitchen. 2021;14:43-54. 10.1000/demo.2021.14.43
12.Andersen M, Maier N. Table header spanning in HTML and print. J Doc Eng. 2022;15:46-57. 10.1000/demo.2022.15.46
13.Maier N, Schmidt O. Structured publishing workflows for small presses. J Scholarly Publ. 2023;16:49-60. 10.1000/demo.2023.16.49
14.Schmidt O, Rossi P. Metadata interoperability in JATS. Inf Stand Q. 2024;17:52-63. 10.1000/demo.2024.17.52
15.Rossi P, Okonkwo Q. Effect sizes in small randomized trials. Trials Today. 2025;18:55-66. 10.1000/demo.2025.18.55
16.Okonkwo Q, Beaumont R. A survey of PDF rendering pipelines. Digit Libr Rev. 2018;19:58-69. 10.1000/demo.2018.19.58
17.Beaumont R, Lindqvist S. Affiliation disambiguation with ROR. Scholarly Kitchen. 2019;3:61-72. 10.1000/demo.2019.3.61
18.Lindqvist S, Hassan T. Hanging indents considered carefully. J Doc Eng. 2020;4:64-75. 10.1000/demo.2020.4.64
19.Hassan T, Nakamura U. Two-column flow in paged media. J Scholarly Publ. 2021;5:67-78. 10.1000/demo.2021.5.67
20.Nakamura U, Petrov V. Footnote placement strategies. Inf Stand Q. 2022;6:70-81. 10.1000/demo.2022.6.70
21.Petrov V, Ferreira W. MathML preservation across converters. Trials Today. 2023;7:73-84. 10.1000/demo.2023.7.73
22.Ferreira W, Kowalski X. Anchor stability in long documents. Digit Libr Rev. 2024;8:76-87. 10.1000/demo.2024.8.76

Notes

1 This clarification expands on a point that would otherwise interrupt the main argument.
2 The distinction here is subtle but matters for reproducibility.
3 See the supplementary material for the full derivation.
4 A reviewer suggested this alternative interpretation, which we find plausible.
5 The exact threshold is implementation-defined and may vary between versions.
6 Historical note: earlier drafts used a different convention.
7 This edge case did not occur in our corpus but is documented for completeness.
8 The terminology follows the convention established in the cited standard.
9 This clarification expands on a point that would otherwise interrupt the main argument.
10 The distinction here is subtle but matters for reproducibility.
11 See the supplementary material for the full derivation.
12 A reviewer suggested this alternative interpretation, which we find plausible.
13 The exact threshold is implementation-defined and may vary between versions.
14 Historical note: earlier drafts used a different convention.
15 This edge case did not occur in our corpus but is documented for completeness.
16 The terminology follows the convention established in the cited standard.
17 This clarification expands on a point that would otherwise interrupt the main argument.
18 The distinction here is subtle but matters for reproducibility.
19 See the supplementary material for the full derivation.
20 A reviewer suggested this alternative interpretation, which we find plausible.
21 The exact threshold is implementation-defined and may vary between versions.
22 Historical note: earlier drafts used a different convention.