Aligning Relative Sequences

library(eratosthenes)

Introduction

This vignette outlines functions that check and align multiple relative sequences of events using eratosthenes. Not all sequences are however of the same informational value. Some sequences may be based on ideal or optimal theoretical assumptions, as with frequency or contextual seriation, and others they may be based on physical relationships, as with soil stratigraphy. It is clear that such theoretical sequences should yield to those based on physical relationships, but this raises the problem of how to coerce one or more theoretical sequences which may contain events that also have physical relationships. For example, one can produce an ideal seriation of contexts that includes discrete, single deposits on the one hand (say a number of separate pits with no overlap) and on the other stratified deposits. In such an optimal seriation, it is possible for those stratified deposits to be seriated “out of order” with respect to their stratified sequence (i.e., the stratigraphy may have been perturbed at one more moments so that their finds assemblages are not well stratified). Accordingly, it is desirable to take the optimal order achieved via seriation and constrain it back to agree with the soil stratigraphy.

Checking Sequence Agreement

Events are structured within two classes of objects, produced by the following functions: events() comprise a single, ordered sequence of events, which must be unique character elements, while sequences() comprise a list of events objects, which must not conflict with one another in their ordering across all events objects. Basically, events() is a more restrictive class of c(), while sequences() is a more restrictive class of list(). The operation of the seq_check() function is run automatically when creating a sequences() object, which will throw an error if any events objects disagree in their ordering.

Both sequences x and y contain the same events in the same order:

x <- events("A", "B", "C", "D", "E")
y <- events("B", "D", "F", "E")
a <- sequences(x, y)

Merging Sequences

The synth_rank() function will use recursion in order to produce a single, “merged” or “synthesized” sequence from two or more sequences. This is accomplished by counting the total number of elements after running a recursive trace through all partial sequences (via the quae_postea() function, on which see below). If partial sequences are inconsistent in their rankings, a NULL value is returned.

x <- events("A", "B", "C", "D", "H", "E")
y <- events("B", "D", "F", "G", "E")
a <- sequences(x, y)
synth_rank(a)
#> Events object of 8 elements:
#>    A, B, C, D, F, H, G, E

Producing a single merged or synthesized sequence is a matter of procedural convenience for the gibbs_ad() function (see the vignette on Gibbs Sampling for Archaeological Dates). To be sure, events missing from one sequence or another could occur at different points in the merged sequence. In the example above, "H" could occur at any point after "D" and before "E", but the synth_rank() function has situated it in between "F" and "G".

If sequences disagree, a NULL value is returned for the synth_rank() function.

Adjusting Sequences

As mentioned above, one may have sequences which are derived via theoretical considerations (e.g., frequency or contextual seriation), and some which are known (e.g., soil stratigraphy or historical documentation). The seq_adj() function will take an “input” sequence and adjust its ordering to fit with another “target” sequence of smaller size.

For example, the input sequence might be an ordering obtained from a contextual seriation which is based on artifact types found in both tomb assemblages as well as stratified deposits, from one or more sites. And the target sequence might be a known stratigraphic sequence. One wants to maintain the ordering of the target sequence, adjusting the input into agreement:

# input
seriated <- events("S1", "T1", "T2", "S3", "S4", "T4", "T5", "T6", "S5", "T7", "S2")
# target
stratigraphic <- events("S1", "S2", "S3", "S4", "S5")
# input adjusted to agree with the target
seq_adj(seriated, stratigraphic)
#> Events object of 11 elements:
#>    S1, T1, S2, T2, S3, T7, S4, T4, T5, T6, S5

To achieve this the seq_adj() function performs a linear interpolation between jointly attested events, placing the input sequence along the \(x\) axis and the target sequence along the \(y\) axis, coercing the order of all elements to the \(y\) axis:

For multiple stratigraphic sequences, the seq_adj() function can be re-run, taking a new target sequence and using the previous result as the new input sequence. If both input and target sequences agree, the input sequence will be returned.

All Earlier Events, All Later Events

A core need of the gibbs_ad() function is to determine, for each event, all events which come later and earlier than that event. The functions quae_postea() and quae_antea() achieve this need for later events and earlier events respectively. Hence, there is no need to determine only one, single sequence or ordering as an input. Instead, a list object which contains multiple (partial or incomplete) sequences. The output is a list indexed with each element, containing the vector of contexts which precede or follow that element.

For quae_antea(), a dummy element of "alpha" is included in all vectors, and for quae_postea(), a dummy element of "omega" is included. The elements of "alpha" and "omega" are necesary as they constitute the fixed lower and upper limits in which estimates are made in gibbs_ad().

x <- events("A", "B", "C", "D", "H", "E")
y <- events("B", "D", "F", "G", "E")

quae_postea(x, y)
#> $A
#> [1] "B"     "C"     "D"     "H"     "E"     "F"     "G"     "omega"
#> 
#> $B
#> [1] "C"     "D"     "H"     "E"     "F"     "G"     "omega"
#> 
#> $C
#> [1] "D"     "H"     "E"     "F"     "G"     "omega"
#> 
#> $D
#> [1] "H"     "E"     "F"     "G"     "omega"
#> 
#> $H
#> [1] "E"     "omega"
#> 
#> $E
#> [1] "omega"
#> 
#> $F
#> [1] "E"     "G"     "omega"
#> 
#> $G
#> [1] "E"     "omega"
quae_antea(x, y)
#> $A
#> [1] "alpha"
#> 
#> $B
#> [1] "A"     "alpha"
#> 
#> $C
#> [1] "A"     "B"     "alpha"
#> 
#> $D
#> [1] "A"     "B"     "C"     "alpha"
#> 
#> $H
#> [1] "A"     "B"     "C"     "D"     "alpha"
#> 
#> $E
#> [1] "A"     "B"     "C"     "D"     "H"     "F"     "G"     "alpha"
#> 
#> $F
#> [1] "A"     "B"     "C"     "D"     "alpha"
#> 
#> $G
#> [1] "A"     "B"     "C"     "D"     "F"     "alpha"

Using events as input into quae_postea() and quae_antea() automatically check of the agreement of events’ ordering. A sequences object can however also be used as input:

x <- events("A", "B", "C", "D", "H", "E")
y <- events("B", "D", "F", "G", "E")
a <- sequences(x, y)

quae_postea(a)
#> $A
#> [1] "B"     "C"     "D"     "H"     "E"     "F"     "G"     "omega"
#> 
#> $B
#> [1] "C"     "D"     "H"     "E"     "F"     "G"     "omega"
#> 
#> $C
#> [1] "D"     "H"     "E"     "F"     "G"     "omega"
#> 
#> $D
#> [1] "H"     "E"     "F"     "G"     "omega"
#> 
#> $H
#> [1] "E"     "omega"
#> 
#> $E
#> [1] "omega"
#> 
#> $F
#> [1] "E"     "G"     "omega"
#> 
#> $G
#> [1] "E"     "omega"
quae_antea(a)
#> $A
#> [1] "alpha"
#> 
#> $B
#> [1] "A"     "alpha"
#> 
#> $C
#> [1] "A"     "B"     "alpha"
#> 
#> $D
#> [1] "A"     "B"     "C"     "alpha"
#> 
#> $H
#> [1] "A"     "B"     "C"     "D"     "alpha"
#> 
#> $E
#> [1] "A"     "B"     "C"     "D"     "H"     "F"     "G"     "alpha"
#> 
#> $F
#> [1] "A"     "B"     "C"     "D"     "alpha"
#> 
#> $G
#> [1] "A"     "B"     "C"     "D"     "F"     "alpha"