AI Tongue Methodology & Limitations
Missing evidence is displayed as missing. A bare “95% accurate” claim is not acceptable without the task, model, dataset, split, metric and limitations.
Transparency Principle
A methodology page is trustworthy when it makes missing evidence visible. “No audited public metric is currently published” is more credible than a fabricated accuracy percentage.
What Is Public Today?
Unknown model and validation values are displayed as Not Published / Evidence Pending rather than invented.
What the Model May Do—and What It Must Not Do
The literal production pipeline must be verified before it is published as architecture.
Intended Educational Pipeline
- Receive a suitable image.
- Evaluate image quality.
- Detect/segment the supported visible region.
- Extract supported visible features.
- Map supported outputs to traditional constitution-oriented context.
- Return confidence and limitations.
Not Intended
- disease diagnosis or screening
- cancer / diabetes / organ-function diagnosis
- anemia or infection detection
- prescription or medication recommendations
- emergency triage
Input Quality Is Part of Performance
Lighting, exposure, white balance, blur, camera processing, distance, tongue position, obstruction and cropping can affect an image-based system.
Acceptable
Proceed under the current tested capture specification.
Borderline
Warn or lower confidence only if the current method supports that state.
Unacceptable
Reject rather than invent a result.
Report Rejection Rate
Not Published
Images Are Not People
Unique people, original images, usable images and augmented images must be counted separately.
Labels, Duplicates & Train/Test Separation
Performance can be overstated if the label definition is weak or the same person's images leak across splits.
Define the Label
Document who assigned it, framework source, annotator process, adjudication and uncertain-label handling.
Control Repeated Images / People
Exact and near duplicates plus same-person grouping matter when person-level generalization is the intended test.
Separate Development from Evaluation
Document splitting unit, proportions/counts, leakage prevention and tuning policy.
A Single “Accuracy” Number Is Not Enough
Every performance statement should name the task, model version, dataset, unit of analysis, split, metric, threshold, subgroup and limitations.
Where the Model Can Fail
Failure conditions should be public and versioned.
Lighting / Blur / Crop
Colored lighting, blur, partial tongue and heavy filtering are obvious failure candidates.
New Cameras / Populations
Performance can shift across country, device and capture workflow.
Ambiguity / Unsupported Cases
Coating/color ambiguity, unusual presentation and out-of-distribution inputs need explicit handling.
Model Changes Require New Evidence
Do not silently retrain and keep the same methodology page/version.
Traditional Framework ≠ AI Performance ≠ Medical Diagnostic Validity
P071 defines constitution concepts. P074 explains the narrow digital model task. Neither one should be used to manufacture medical diagnostic validity.
Frequently Asked Questions
Short answers preserve the page boundary and route deeper safety or product questions to the correct owner.
How accurate is AI tongue analysis?
Use the published model/task metrics. If no audited metrics exist, no public validated percentage should be stated.
How many images trained the AI?
See audited original and usable image counts when published.
How many people?
See the separate unique-person count.
Has it been externally validated?
Use the explicit external-validation status.
Does lighting affect results?
It can; current capture and robustness tests should quantify the effect.
Is confidence the same as accuracy?
No.
Is this a disease diagnostic model?
No.
Can an advisor replace validation?
No.
Can a model update change accuracy?
Yes; material updates require new validation and version disclosure.