Abstract
The question of how to evaluate the epistemic quality of an unpublished literary piece/manuscript which may be a working paper, a conference paper, a pre-institutional draft circulating in scholarly networks, an affidavit (unsigned) etc— has remained mostly unaddressed in the existing literature on AI training data governance. Importantly, this is not a small concern. The PWTA, published by De and Pal,1 suggests the scoring of published legal sources with considerable rigor. However, for an unpublished document, however, the four-tier institutional hierarchy of this framework, which presupposes that every source has cleared some form of institutional validation, cannot be applied. The Āpata Filter1 assigns such a document a machine training weight of zero, which is epistemically correct. It does not provide is a human-facing audit score for the manuscript’s intrinsic quality, grounded in the same mathematically documented, unbiased methodology that the rest of the PWTA architecture employs.
Here, we address this gap by the development of Pramana-Weighted Training Architecture Unpublished Document Extension, designated as PWTA-UDE. The extension operationalizes the classical BhāṭṭaMīmāṃsā doctrine of anupalabdhi which is valid cognition through the noticed and correctly identified absence of expected evidence. The system is developed into a measurable, documented Transparency Score (TS_unpub) for pre-institutional documents. The four statistically scoring components have been developed: the Bibliographic Epistemic Lineage (R_EPS), the Epistemic Sufficiency Score (E_s), the Neo-Empirical Content Score (C_neo), and the Syntactic Write Quality Score (S_write). An “Anupalabdhi”-derived fixed discount factor of 0.85 is applied multiplicatively in the final step, establishing a structural ceiling at TS_unpub ≤ 0.85 for all unpublished documents.
This study reports the results of applying PWTA-UDE to five separate manuscripts from different law genres. The resulting scores show a sharp gradient of 0.4821 for the data-rich academic manuscript, 0.1501 for the professionally drafted petition, and 0.0000 for every form of commentary. The system draws the discrimination that the extension was designed to make. It was made without any element of human judgment about any author. For every document, the Āpata lock held PWTA(s) = 0.0000 independently. The framework is further examined in the light of comparative administrative (regulative) developments, such as those in the, EU, the United-Kingdom, and the United-States of America, and the UNESCO and NITI Aayog ethics standards.
Keywords: Pramāṇa-Weighted Training Architecture, PWTA-UDE, Anupalabdhi, Bhāṭṭa Mīmāṃsā, Epistemic Scoring, Legal AI Training Data Governance, Indian Knowledge Systems
- Introduction
Artificial intelligence systems are currently deployed in Indian courts, government departments, and administrative bodies. The Supreme Court of India’s SUPACE platform performs legal research and in some cases analysis also. State welfare departments use automated eligibility screening. Computationally, an AI system’s output remains a direct function of the quality of its training corpus.2 A system trained on a general mixture of primary legislation, anonymous commentary, and commercially motivated content learns from all of them together, with no internal mechanism to distinguish between and no record of the relative weight each was given.
As Indian courts and administrative bodies rapidly adopt AI system platforms and AI based software, the constitutional dimension of training data quality has become well-established. The non-arbitrariness standard in Maneka Gandhi v. Union of India3which mandates that AI-assisted official conduct rest on principled, evidential foundations. The Digital Personal Data Protection Act, 20234 imposes data accuracy obligations on data fiduciaries/ trustees. The professional liability is extended so that it may reach AI-assisted decisions in Bharatiya Nyaya Sanhita, 20235. Also, the IndiaAI Mission6across the government deployments it mandates a responsible AI development. The NITI Aayog’s Responsible AI for All7takes into account transparency, anti bias and fairness as the core virtues. Still, none of these instruments prescribes a computable standard for evaluating the epistemic quality of pre-institutional training sources. This is the concern which this study tries to identify.
The Pramāṇa-Weighted Training Architecture, published by De and Pal,1 suggests a rigorous, four-tier, statistical framework for scoring published legal sources through epistemic study and mathematical applications. The framework assigns machine training weights based on a source’s institutional tier and content quality, derived from a real calibration corpus and applied through closed-form formulae.
However, for an unpublished manuscript a demonstrable lacuna exists. That is, the four-tier hierarchy presupposes institutional validation, which an unpublished document by definition has not cleared. Therefore, the Āpata Filter1 assigns it a training weight of zero. This outcome is architecturally correct and must not be disturbed. Consequently, a parallel, human-facing score that evaluates the document’s intrinsic quality through the same methodology is absent.
To bridge the above-mentioned lacuna, we propose PWTA-UDE, which is designed as a single-pass extension generating a transparency Score for pre-institutional manuscripts. Although the extension preserves the Āpata lock and introduces no human judgment into the scoring process, it is further calibrated against a real corpus of 177 Indian legal academic documents. It has been applied to five different manuscripts of different writers and belonging to separate law genres calculating every metric directly from raw textual data derived from direct measurement of that paper’s actual text to demonstrate its applicability.
The paper is structured as follows. Section II reviews the existing literature. The Section III of this study helps to identify the specific research gap. Subsequently, Section IV describes the PWTA-UDE framework then the Section V details the mathematical tools to be employed. The Section VI applies PWTA-UDE and presents the results obtained from real measurements, the Section VII then examines its applications in the field of law. Lastly, Section VIII provides the conclusion with suggestions.