This independent open-research project evaluates possible additions to the Unicode Standard. It develops evidence both for and against encoding prospective characters and repertoires; inclusion in the repository does not confer official candidate status or imply that a proposal should be submitted.
The project concentrates on prospective, generally unencoded characters and symbols that may have a case for encoding. Physical signs and signage are studied as evidence of real-world symbol use, not as the entities Unicode would encode.
Some candidates may be viable through the character-proposal process. Others are emoji-first and may be viable only through the emoji-proposal process; U+1F965 🥥 COCONUT is a useful precedent. A candidate may also warrant investigation through both routes. Emoji characters and emoji sequences are plain text; the distinction here concerns proposal routes, not text versus emoji.
The repository covers:
- individual prospective characters;
- repertoires whose characters must be researched together;
- shared studies that inform several candidate projects;
- reusable research and proposal methods; and
- versioned submission material when a proposal reaches that stage.
Standalone proposals concerning CLDR data, Unicode Standard Annexes, character properties, or other non-character technical changes are outside the project’s present scope.
The candidate inventory is maintained separately so the repository’s primary view explains the research purpose before presenting subjects under evaluation.
| Location | Purpose |
|---|---|
| Candidates | Research specific to candidates with contributed material |
| Research | Substantive studies that apply across candidates |
| Methods | Reusable workflows and refinement logs |
| Project | Project-wide lifecycle and governance decisions |
| Scripts | Lightweight project-record validation |
Folders are added only when there is material to place in them. Related but independently viable symbols retain separate candidate projects even when they share research or may later appear in one proposal.
Research distinguishes observed evidence, sourced claims, working inferences, and proposal decisions. It records material provenance, source quality, access dates, geographic and usage context, and publication rights where available. Negative and unresolved findings are retained when they improve the assessment.
Research received from Gemini, NotebookLM, ChatGPT, or another external model follows External model research intake: model output is treated as a lead or analytical artefact, and the underlying source must be checked before a finding enters integrated project research.
Original evidence is preserved unchanged. Derivatives such as crops or annotations remain linked to their originals. Availability, exposure, observed symbol use, and demonstrated comprehension are treated as separate evidential steps.
Viewing, discussion, corrections, evidence contributions, and methodological collaboration are welcome. The project is being opened before a mature contribution framework is designed; early contributions should identify their source, provenance, and reuse rights as clearly as possible.
Original project text and assets are available under the terms described in LICENCE.md. Third-party material is excluded from that blanket licence and retains its own terms.
Unicode® is a registered trademark of Unicode, Inc. in the United States and other countries. This independent project is not associated with, endorsed by, or sponsored by Unicode, Inc.