Skip to content

§03-ii Minerva — Knowledge Base — Inputs and corpus characterization

Mars® Spec§03-ii Minerva — Knowledge Base › Inputs and corpus characterization

← What a knowledge base is and what distinguishes it · Section index · §3 — Aspect-typed primitive extraction (Nucleus A) →

2. Inputs and corpus characterization

Corpus input. The primary input is a corpus of unstructured data. Text is the primary form: documents, books, papers, reports, transcripts, code, or any text-bearing content. Non-language data — figures, images, audio, tabular records, structured records — requires a language-bound interpretation before it can be chunked and processed: the data is bound to a text representation that expresses its meaning in a form amenable to chunking, embedding, and aspect-typed extraction. The language-bound interpretation is not a translation or summary; it is a governed representation that preserves the semantic content of the non-language data in a form the KB can store and retrieve. The native-form data and its language-bound interpretation are both retained: the native form in the non-embedded store (or its native storage, referenced by the KB), and the interpretation as the chunkable and embeddable representation.

Citation and source location provenance. Every corpus item carries citation and source location provenance at the point of ingestion: the source identity (document, dataset, or collection), the location within the source (page, section, paragraph, timestamp, or equivalent locator), and any available bibliographic metadata (author, date, publisher, identifier). This provenance is attached to every chunk and subchunk derived from the item and is carried through to every extracted primitive. An inspector can trace any KB entry back to its corpus origin via this provenance chain.

Optional inputs. One or more existing knowledge bases may be provided. Where provided, they participate in extraction, gap-filling, and merge operations (§3.4).

Model/KB relationship. The domain model and the knowledge base are mutually generative. A corpus may be used in corpus-seed mode to generate an order-decomposed model (§01 §3); a mature KB may participate at any downstream stage including seed derivation, semantic binding, query-library construction, and gap-filling. No hierarchy is imposed between model and KB.

KB metadata is a property of the KB, built and versioned against the governing domain model version. Every metadata entry in a knowledge base — aspect identifiers, primitive bindings, chunk and subchunk identifiers, embedding vector identifiers, provenance records, and configuration artifacts — is built relative to a specific domain model version. The domain model’s order-decomposition tree determines the aspect identifiers; the domain model’s version determines the scope of every extraction, binding, and synthesis act. A KB’s metadata is not domain-agnostic: it is an expression of the governing domain model version applied to the corpus. Where multiple domain models are active simultaneously, each is composable and independently operable (§01 §1); a KB built against one domain model version carries metadata scoped to that version, independently of any other active domain model. When a domain model version is amended, the KB entries built against that version are subject to the staleness and recomputation obligations of §3.4; entries built against other active domain model versions are not affected. Two KBs built against different domain model versions carry structurally distinct metadata and are not directly comparable without a model-version reconciliation step.

Bootstrapping the domain model from KB(s) and externally-provided specifications. When a governing domain model is absent, it can be bootstrapped from one or more knowledge bases or externally-provided specifications available as inputs. An externally-provided specification — a protocol, standard, guideline, rules document, or other externally-authoritative artifact — is source material that can be ingested as corpus for bootstrapping; it is not itself a formal specification before this process, because a formal specification is the lowered output of a domain model and is produced by predicate lowering, not an input to one. Bootstrapping proceeds by applying the corpus-seed mode of §01 §3 to the available KB(s) or externally-provided specification(s) treated as corpus, producing an order-decomposed model that becomes the governing reference against which subsequent KB metadata is built. A bootstrapped domain model is a first-class registered artifact enrolled in the delta-attestation lifecycle (§02b §5); it is not a provisional or temporary artifact. Once registered, it governs KB metadata construction identically to any other domain model, and the formal specification (the lowered output of that domain model) can then be produced through predicate lowering. Reversing a formal specification back into the domain model that generated it recovers at best the same model and is not a bootstrap path.


1.1 Formal specification as KB variant

A formal specification is the lowered output of a domain model, produced by predicate lowering: the act of applying the domain model’s order structure, aspects, and perspectives to available knowledge — rules, evidence, corpus content, general knowledge — to produce governed, order-typed assertions about the domain. The formalization act — predicate lowering through the domain model — is what distinguishes a formal specification from a general-purpose KB. A formal specification is always downstream of a domain model; it cannot be an input to the domain model that generated it (that would recover at best the same model), and it cannot be produced without a domain model.

Direction note. An externally-provided specification — a regulatory document, technical standard, contractual terms document, or domain authority publication — is source material that can bootstrap a domain model (§01 §2.8) or serve as corpus for KB construction. It is not a formal specification before that process. Once a domain model has been established (whether bootstrapped from the externally-provided specification or derived independently), the formal specification is produced by predicate lowering through that domain model. Declaring an externally-provided document as a formal specification at the point of registration without predicate lowering through a domain model is not a conforming path to a formal specification under this section.

What source material can feed the formalization act. The formalization act — predicate lowering through the domain model — may draw from any combination of available knowledge; there is no restriction on the source material fed into that act. Source material may include:

  • Extracted corpus content — derived from a corpus through the KB extraction pipeline (§3.1), where the extraction is governed by the domain model and the extracted content is lowered into order-typed assertions about the domain.
  • Built-up knowledge — assembled from rules, evidence, and general knowledge through any combination of governed operations, including higher-order synthesis (§01 §11), governed data generation (§03-iv §2), and iterative reinforcement; the result is lowered through domain mapping into order-typed assertions.
  • Externally-provided specification content — a protocol, standard, guideline, or rules document ingested as corpus and lowered through the domain model into order-typed assertions. The source document is the input to the KB extraction pipeline; the formal specification is the output of the predicate lowering act over that content.

The enumeration above is non-limiting. Any source material may feed the formalization act. What produces a formal specification is not the source but the predicate lowering act through the governing domain model.

Formalization as the governing act. Formalization is the act of applying the available knowledge — rules, evidence, corpus content, general knowledge — to the domain model’s order structure, aspects, and perspectives to produce assertions that are: (a) order-typed against the governing domain model, (b) aspect-bound to the domain model’s aspect structure, and (c) registered as a versioned artifact enrolled in the delta-attestation lifecycle. An artifact that carries content relevant to a domain but has not been formalized through domain mapping is not a formal specification; it is source material for a KB or for a formalization act.

A formal specification is a KB variant, not a separate artifact class. Its construction pipeline, metadata structure, co-indexed bidirectional store, and lifecycle are those of the KB (§3.1–§3.4). What distinguishes it from a general-purpose KB is the additional formalization act and the adjudication gate that makes it authoritative as a governing reference. A registered formal specification is a five-component bundle (§01 §7): compiled predicates, well-formedness discharge record, adjudication record, scope-of-applicability certificate, and soundness declaration. All five components are co-required; a bundle missing any component is refused at registration.

Validity determination and adjudication. Before a formal specification may be used as a governing reference — as a constituent of a governing specification (§03-iii §2.2) — it must carry a certifying adjudication record produced by the adjudication gate (§01 §8). The adjudication record is the third component of the formal specification bundle; it is the same artifact this section refers to as validity determination. There is one gate and one record — the adjudication gate at §01 §8 — viewed from two directions: from the modeling layer as boundary-closure attestation, and from this layer as the governing-reference admission check. A formal specification whose adjudication record carries a conditional trust-signal class (latent adjudication, pending full certification) may be used as source material and as a KB input but may not serve as a governing reference until its adjudication record is upgraded to certifying status. A formal specification carrying a certifying adjudication record — human or machine-assisted human adjudication with κ at or above the registered threshold, named reviewers with consequence-bearing standing — is admissible as a governing reference. The adjudicator set for a formal specification is a registered artifact enrolled in the delta-attestation lifecycle. A formal specification may have one or more adjudicators; where the spec is composite — drawing from multiple disciplines, jurisdictions, or bodies of expertise — multiple adjudicators are expected, each scoped to their domain of authority. The same authority-procedure machinery as §03-iii §2.0 governs the adjudication act.

Amendment lifecycle. Amendments to a formal specification are governed by the delta-attestation lifecycle (§02b §5). A scope-breaking amendment invalidates the validity determination; re-adjudication is required before the amended specification may serve as a governing reference. Scope-preserving and scope-extending amendments preserve the validity determination; the adjudicator set is notified of the amendment and may initiate re-adjudication at its discretion.



← What a knowledge base is and what distinguishes it · Section index · §3 — Aspect-typed primitive extraction (Nucleus A) →