Skip to content

Repository files navigation

Unicode Candidate Research

This independent open-research project evaluates possible additions to the Unicode Standard. It develops evidence both for and against encoding prospective characters and repertoires; inclusion in the repository does not confer official candidate status or imply that a proposal should be submitted.

The project concentrates on prospective, generally unencoded characters and symbols that may have a case for encoding. Physical signs and signage are studied as evidence of real-world symbol use, not as the entities Unicode would encode.

Some candidates may be viable through the character-proposal process. Others are emoji-first and may be viable only through the emoji-proposal process; U+1F965 🥥 COCONUT is a useful precedent. A candidate may also warrant investigation through both routes. Emoji characters and emoji sequences are plain text; the distinction here concerns proposal routes, not text versus emoji.

Scope

The repository covers:

  • individual prospective characters;
  • repertoires whose characters must be researched together;
  • shared studies that inform several candidate projects;
  • reusable research and proposal methods; and
  • versioned submission material when a proposal reaches that stage.

Standalone proposals concerning CLDR data, Unicode Standard Annexes, character properties, or other non-character technical changes are outside the project’s present scope.

The candidate inventory is maintained separately so the repository’s primary view explains the research purpose before presenting subjects under evaluation.

Organization

Location Purpose
Candidates Research specific to candidates with contributed material
Research Substantive studies that apply across candidates
Methods Reusable workflows and refinement logs
Project Project-wide lifecycle and governance decisions
Scripts Lightweight project-record validation

Folders are added only when there is material to place in them. Related but independently viable symbols retain separate candidate projects even when they share research or may later appear in one proposal.

Research approach

Research distinguishes observed evidence, sourced claims, working inferences, and proposal decisions. It records material provenance, source quality, access dates, geographic and usage context, and publication rights where available. Negative and unresolved findings are retained when they improve the assessment.

Research received from Gemini, NotebookLM, ChatGPT, or another external model follows External model research intake: model output is treated as a lead or analytical artefact, and the underlying source must be checked before a finding enters integrated project research.

Original evidence is preserved unchanged. Derivatives such as crops or annotations remain linked to their originals. Availability, exposure, observed symbol use, and demonstrated comprehension are treated as separate evidential steps.

Participation

Viewing, discussion, corrections, evidence contributions, and methodological collaboration are welcome. The project is being opened before a mature contribution framework is designed; early contributions should identify their source, provenance, and reuse rights as clearly as possible.

Licence and trademark

Original project text and assets are available under the terms described in LICENCE.md. Third-party material is excluded from that blanket licence and retains its own terms.

Unicode® is a registered trademark of Unicode, Inc. in the United States and other countries. This independent project is not associated with, endorsed by, or sponsored by Unicode, Inc.

About

Open research into prospective Unicode characters and repertoires

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages