Page MenuHomeDevCentral

Support Grapheme functions for UTF-8 strings
ClosedPublic

Authored by dereckson on Feb 21 2022, 23:04.
Tags
None
Referenced Files
F3762956: D2550.id6432.diff
Thu, Nov 21, 18:12
F3762947: D2550.id9264.diff
Thu, Nov 21, 17:57
F3762939: D2550.id9264.diff
Thu, Nov 21, 17:45
F3762692: D2550.diff
Thu, Nov 21, 13:15
Unknown Object (File)
Tue, Nov 19, 15:18
Unknown Object (File)
Tue, Nov 19, 14:28
Unknown Object (File)
Tue, Nov 19, 06:46
Unknown Object (File)
Tue, Nov 19, 06:46
Subscribers
None

Details

Summary

The intl extension supports the grapheme concept from Unicode,
while the mbstring extension handle UT8 codepoints.

An emoji like 🏴󠁧󠁢󠁥󠁮󠁧󠁿 has 28 bytes, 7 codepoints (tags E N G L A N D),
1 grapheme. That could affects method like substr or strlen when
we want to manipulate graphemes and not codepoints.

Strategy is to offer bytes/codepoints/graphemes capabilities,
downgrade from graphemes to codepoints for non UTF-8 encoding,
and defaults to grapheme.

Test Plan

Tests added for new methods.

Diff Detail

Repository
rKERUALD Keruald libraries development repository
Lint
No Lint Coverage
Unit
No Test Coverage
Branch
mbstring-to-grapheme
Build Status
Buildable 3987
Build 4239: arc lint + arc unit

Event Timeline

dereckson held this revision as a draft.

We need unit tests for this change.

Spacing issues. Adding tests.

dereckson retitled this revision from WIP: support Grapheme functions for UTF-8 strings to Support Grapheme functions for UTF-8 strings.Mon, Nov 11, 23:50
dereckson edited the test plan for this revision. (Show Details)
dereckson published this revision for review.Sun, Nov 17, 00:37
dereckson accepted this revision.
This revision is now accepted and ready to land.Sun, Nov 17, 00:37
This revision was landed with ongoing or failed builds.Sun, Nov 17, 00:38
This revision was automatically updated to reflect the committed changes.