Skip to contents

Review representations and projections

Betwixt separates the representation used to prepare a review task from the representations produced after human review. Candidate assertions are prepared in a wide tabular representation that can be created programmatically or imported from ordinary tabular data.

The review task itself can be designed in different ways. In a wide review, evidence supports assertions about a subject. In a dual-wide review, the relation between the evidence and the subject is also brought inside the review boundary.

After human review, the resulting review state can be projected into wide review planes or into long-form atomic semantic assertions.

candidate dataset
       ↓
 review design
   ↙       ↘
wide     dual-wide
   ↘       ↙
  human review
       ↓
 reviewed state
   ↙       ↘
wide       long
projection projection

These distinctions serve different purposes. Wide and dual-wide concern what the reviewer is asked to evaluate. Wide and long projections concern how the resulting review state is represented for subsequent use.

The Dēliņi example

The delini dataset is a small cultural heritage example associated with the Dēliņi farmstead at the Ethnographic Open-Air Museum of Latvia. It contains five observations concerning three heritage artefacts and two documentary records.

data("delini")

delini
#> # A tibble: 5 × 16
#>   row_number evidence_url evidence_media_url     evidence_text label description
#>        <int> <chr>        <chr>                  <chr>         <chr> <chr>      
#> 1          1 NA           https://betwixt.datao… P7101565      Deli… the farmho…
#> 2          2 NA           https://betwixt.datao… P7101561      tabl… a tablet-w…
#> 3          3 NA           https://betwixt.datao… P7101556      bed … a bed in t…
#> 4          4 NA           https://betwixt.datao… P7101590      reco… a record c…
#> 5          5 NA           https://betwixt.datao… P7101623      reco… a floor pl…
#> # ℹ 10 more variables: subject <chr>, subject_range <chr>,
#> #   subject_definition <chr>, instance_of <chr>, instance_of_range <chr>,
#> #   instance_of_definition <chr>, heritage_of <chr>, heritage_of_range <chr>,
#> #   heritage_of_definition <chr>, context_1 <chr>

Each row combines several kinds of information required for a review task:

  • evidence identifies or presents material supporting the observation;
  • description provides human-readable information for the reviewer;
  • candidate columns contain semantic assertions proposed for review;
  • context provides useful information outside the current review boundary.

These roles are represented with ordinary columns rather than a specialised R object. This makes the candidate dataset suitable for exchange through tabular formats such as CSV and Excel.

Candidate columns

A candidate column represents something that can be reviewed. In delini, the principal candidate columns are:

subject
instance_of
heritage_of

A candidate column may be accompanied by two related columns:

<name>
<name>_range
<name>_definition

For example:

instance_of
instance_of_range
instance_of_definition

The candidate column contains the proposed value. Its _range column can provide controlled or suggested alternatives, while its _definition column can provide a semantic definition or reference associated with the candidate.

For the instance_of candidate in delini, these columns distinguish the value proposed for a particular observation from the range available to the reviewer and the semantic resource defining the relation.

In R, these columns can be added programmatically with add_candidate_column(), add_candidate_range() and add_candidate_definition(). The tabular convention itself is independent of R: the same structure can be prepared manually in CSV or Excel and imported as a candidate dataset.

Evidence, description and context

Evidence is deliberately kept distinct from the candidate assertions being reviewed. The Dēliņi example can provide an evidence resource through evidence_url, directly displayable media through evidence_media_url, and a short identification through evidence_text.

The label and description columns help the reviewer understand the resource but are not automatically reviewable assertions.

Similarly, context_1 provides information useful during review without bringing that information inside the current review boundary. In the Dēliņi example it identifies the institution holding the material.

This distinction allows a review task to present rich supporting information without treating everything visible to the reviewer as a candidate assertion.

Wide review

The candidate dataset is deliberately wide. A single row keeps an observation, its evidence and several related candidate assertions together:

Evidence | Subject | instance_of | heritage_of | Context

For domain experts this is often more natural than presenting every semantic assertion as a separate row. The reviewer can inspect the resource once and consider several proposed assertions about it together.

In an ordinary wide review, the evidence supports the review task but its relation to the subject is not itself reviewed.

Evidence ──────► Subject | instance_of | heritage_of
  context           └──────── reviewable ──────────┘

Rendering the wide review

render_review() transforms the candidate dataset into a standalone browser-based review artefact.

wide_review_file <- render_review(
  delini,
  cols = c(
    subject = "Subject",
    instance_of = "instance of",
    heritage_of = "heritage of",
    context_1 = "held by"
  ),
  subheadings = c(
    instance_of = "type",
    heritage_of = "cultural association"
  ),
  title = "Dēliņi semantic review",
  description = "Review the proposed semantic assertions.",
  project_id = "delini",
  filename_stem = "delini-wide",
  sequence = 0L,
  path = tempdir()
)
#> Review rendered: C:\Users\DANIEL~1\AppData\Local\Temp\RtmpO8BEIO/delini-wide.html

Rendering does not turn candidate values into accepted knowledge. It creates a human review environment in which proposed values can be corroborated, corrected, rejected or deferred.

The resulting HTML combines the candidate assertions with the evidence and descriptive information needed to evaluate them. It also records the review state and provenance required to read the review back into Betwixt.

Dual-wide review

Sometimes the relationship between evidence and subject is itself uncertain or semantically important. A photograph may depict an artefact, while an archival record may document it.

A dual-wide review brings this relation inside the review boundary:

Evidence ─relation─► Subject | instance_of | heritage_of
          └────────────── reviewable ──────────────┘

For the Dēliņi example, we can create a second candidate dataset in which the evidence relation is explicitly reviewable.

delini_dual <- create_candidate_dataset(
  evidence_url = delini$evidence_url,
  evidence_media_url = delini$evidence_media_url,
  evidence_text = delini$evidence_text,
  label = delini$label,
  description = delini$description,
  subject = delini$subject,
  subject_range = delini$subject_range,
  subject_definition = delini$subject_definition,
  evidence_relation = c(
    "depicts", "depicts", "depicts", "documents", "documents"
  ),
  evidence_relation_range = add_candidate_range(
    "depicts", "documents", "Other…"
  )
)

delini_dual <- delini_dual |>
  add_candidate_column(
    name = "instance_of",
    value = delini$instance_of,
    range = delini$instance_of_range,
    definition = delini$instance_of_definition
  ) |>
  add_candidate_column(
    name = "heritage_of",
    value = delini$heritage_of,
    range = delini$heritage_of_range,
    definition = delini$heritage_of_definition
  )

The two candidate datasets contain the same assertions about the heritage resources. They differ in the boundary of the review task.

In the dual-wide representation, evidence_relation is itself a candidate assertion. It remains distinct from review provenance, which records who performed the review, when the review took place, and what decisions were made.

Rendering the dual-wide review

The same renderer is used for the dual-wide candidate dataset:

dual_review_file <- render_review(
  delini_dual,
  cols = c(
    evidence_relation = "Evidence relation",
    subject = "Subject",
    instance_of = "instance of",
    heritage_of = "heritage of",
    context_1 = "held by"
  ),
  title = "Dēliņi dual-wide semantic review",
  description = "Review the evidence relation and proposed assertions.",
  project_id = "delini",
  filename_stem = "delini-dual-wide",
  sequence = 0L,
  path = tempdir()
)
#> Review rendered: C:\Users\DANIEL~1\AppData\Local\Temp\RtmpO8BEIO/delini-dual-wide.html

Wide and dual-wide are therefore not different semantic models. They are different review designs. The wide review treats evidence as supporting material. The dual-wide review additionally asks the reviewer to evaluate how that evidence relates to the subject.

From review to reviewed state

After a review has been saved, read_review() reads the standalone review artefact back into Betwixt.

review <- read_review("delini-wide-finalised.html")

The resulting review object preserves the distinction between what was originally proposed and what resulted from human review.

The reviewed state can subsequently be projected in two useful forms:

                    ┌─ wide → candidate / reviewed / status
reviewed state ─────┤
                    └─ long → atomic semantic assertions

These are projections of the same reviewed state rather than alternative review layouts.

Wide review projection

project_review_wide() exposes three related planes of the review:

wide <- project_review_wide(review)

The candidate plane records the values originally proposed for review. The reviewed plane records the values resulting from human review. The status plane records the review outcome associated with each reviewable value.

Conceptually:

candidate   proposed value
reviewed    value after human review
status      review outcome

Keeping these planes distinct makes changes introduced through review explicit rather than overwriting the original candidate material.

Long assertion projection

project_review_long() projects the reviewed state into atomic semantic assertions:

long <- project_review_long(review)

Instead of keeping several candidate predicates on one wide row, the long projection represents individual assertions explicitly:

subject | predicate | value | status

A wide review row can therefore produce several atomic assertions. The long representation makes the subject–predicate–value structure explicit and is suited to subsequent semantic processing and serialisation.

The long projection does not introduce new knowledge. It projects the reviewed state into a representation in which the individual assertions and their review status are explicit.

Review design and semantic projection

Betwixt therefore distinguishes two operations that can otherwise look superficially similar.

Review design determines what is presented to the reviewer:

wide       evidence supports the reviewed assertions
dual-wide  evidence relation is itself reviewable

Projection determines how the resulting review state is represented:

wide       candidate, reviewed and status planes
long       atomic subject–predicate–value assertions

This separation allows the review interface to remain suited to human judgement while the resulting reviewed knowledge can be projected into forms suited to subsequent computational use.

candidate assertions
        ↓
    review design
        ↓
    human review
        ↓
   reviewed state
        ↓
semantic projection
        ↓
subsequent processing
and serialisation