TokenVector

English | Tiếng Việt

TOKENVECTOR COMPILER PLATFORM - RELEASE PACKAGE & TECHNICAL REPORT

(Official Release Package & Comprehensive 3-Pole Technical Benchmark Report)

Document Code: TKV-RELEASE-2026-MASTER
Version: 2026.1 (Independent Commercial Release)
Copyright & Validation: TokenVector Compiler Engineering Team & Antigravity AI Team
GitHub Repository: TokenVector Repository
GitHub Pages Website: TokenVector Project Page


⭐️ Support the Project

If TokenVector empowers your development workflow, accelerates your execution speed, or provides a valuable alternative for high-performance computing, please consider giving it a Star on GitHub!
Your support boosts the project’s visibility and fuels the ongoing development of the TokenVector compiler platform and its ecosystem.

GitHub stars


📦 I. RELEASE DIRECTORY HIERARCHY (release/)

In compliance with professional software release management standards, all deliverables are centrally organized under the release/ root directory:


⚡ II. TOP 5 HIGHLIGHT STATS


📊 III. 5-STAR RATING MATRIX

Evaluation Criteria CPython 3.12 TokenVector (AOT Binary) C++ Native
Ease of Development ⭐⭐⭐⭐⭐ (Easiest) ⭐⭐⭐⭐⭐ (100% Python Syntax) ⭐⭐ (Complex, manual pointer management)
Packaging & Distribution ⭐⭐ (Requires venv / interpreter) ⭐⭐⭐⭐⭐ (Standalone .exe, ~8.5-9 KB measured) ⭐⭐⭐⭐⭐ (Standalone native binary)
Multicore Concurrency ⭐ (Throttled by GIL) ⭐⭐⭐⭐⭐ (No GIL, ~25.9x faster measured on int workload) ⭐⭐⭐⭐⭐ (Full parallel throughput)
Single-Thread Compute ⭐⭐⭐ (Bytecode interpretation) ⭐⭐⭐⭐ (AOT CIL Unboxed Native) ⭐⭐⭐⭐⭐ (Native Machine Code Compilation)
Library Ecosystem ⭐⭐⭐⭐⭐ (PyPI 500k+ pkgs) ⭐⭐⭐⭐ (Python + C-FFI + .NET BCL) ⭐⭐⭐⭐ (C/C++ Ecosystem)

🚀 IV. PURE IN-PROCESS ALGORITHM EXECUTION SPEED BENCHMARK

Empirically measured on 2026-08-31 (median of 3 runs, identical hardware host, self-hosted tkvc.exe vs CPython 3.12.10). C++ column omitted: the benchmarking environment lacked g++/cl.exe toolchains to verify — prior legacy C++ metrics were discarded to prevent unsubstantiated claims.

(The FP64 multithreaded test case initially exposed a real compiler bug — thread_join() performed an invalid type cast when workers returned f64, triggering a runtime InvalidCastException — which was patched on the same day; refer to docs/BUGS_TODO.md.)

Benchmark Test (Workload) CPython 3.12 (median) TokenVector AOT (median) Ratio
Integer Loop (10M Ops) 1,852 ms 82 ms TokenVector is 22.6x faster than Python
Floating-Point Arithmetic FP64 (2M Ops) 290 ms 18 ms TokenVector is 16.1x faster than Python
Multithreaded Integer (4 Threads x 5M) 3,284 ms 127 ms TokenVector is 25.9x faster than Python (No GIL)
Multithreaded Float (4 Threads x 2M Float) 1,222 ms 65 ms TokenVector is 18.8x faster than Python (No GIL)

💾 V. BINARY FOOTPRINT, COMPILATION LATENCY & PACKAGING COMPARISON

Empirically measured on 2026-08-31. C++ column omitted (for rationale, see Section IV).

Technical Metric CPython 3.12 TokenVector AOT PE
Packaged Distribution Size (Compiled Program .exe) 25 MB - 100 MB (Runtime dependent) ~8.5 - 9 KB (Standalone, empirically measured)
Compilation Latency (Build Time via tkvc.exe) 0 ms (Instant bytecode generation) ~2.3 - 3.9 seconds (empirically measured) — largely PyInstaller-frozen bootstrap overhead of tkvc.exe (the compiler is written in .tkv but currently packaged via CPython freezing, NOT self-compiled to native executable), not the internal AST→IL compilation logic
External Environment Dependencies Strict requirement for Python runtime + DLLs ZERO CPython Requirement (compiled .tkv programs run standalone; tkvc.exe binary builder itself has dependencies, noted above)
Intellectual Property Protection (Reverse Eng) Easily decompiled back to source (.pyc) Compiled outputs protected by AOT CIL Assembly

💻 VI. 3-WAY CODE COMPARISON MATRIX (TOKENVECTOR VS PYTHON VS C++)

⚡ Column 1: TokenVector (Untitled-1.tkv)

# -*- coding: utf-8 -*-
class DataAnalyzer:
    name: "str"
    baseline: "f64"
    def __init__(self, name, baseline):
        self.name = name
        self.baseline = baseline

def compute_performance(name: "str", baseline: "f64", score1: "f64", score2: "f64") -> "f64":
    analyzer = DataAnalyzer(name, baseline)
    avg = (score1 + score2) / 2.0
    return avg - analyzer.baseline

def process_numbers(limit: "i32") -> "i32":
    sum_val = 0
    for i in range(1, limit + 1):
        sum_val = sum_val + i
    return sum_val

def main() -> "i32":
    print("=== TOKENVECTOR NATIVE ===")
    delta = compute_performance("Core", 50.0, 85.0, 95.0)
    total_sum = process_numbers(100)
    print("Delta: " + str(delta))
    print("Sum: " + str(total_sum))
    return 1

🐍 Column 2: Python 3 (Untitled-1.py)

# -*- coding: utf-8 -*-
class DataAnalyzer:
    def __init__(self, name: str, baseline: float):
        self.name = name
        self.baseline = baseline

def compute_performance(name: str, baseline: float, score1: float, score2: float) -> float:
    analyzer = DataAnalyzer(name, baseline)
    avg = (score1 + score2) / 2.0
    return avg - analyzer.baseline

def process_numbers(limit: int) -> int:
    sum_val = 0
    for i in range(1, limit + 1):
        sum_val = sum_val + i
    return sum_val

def main() -> int:
    print("=== PYTHON CPYTHON ===")
    delta = compute_performance("Core", 50.0, 85.0, 95.0)
    total_sum = process_numbers(100)
    print("Delta: " + str(delta))
    print("Sum: " + str(total_sum))
    return 1

⚡ Column 3: C++20 (Untitled-1.cpp)

#include <iostream>
#include <string>

class DataAnalyzer {
public:
    std::string name;
    double baseline;
    DataAnalyzer(std::string n, double b) : name(n), baseline(b) {}
};

double compute_performance(std::string name, double baseline, double score1, double score2) {
    DataAnalyzer analyzer(name, baseline);
    double avg = (score1 + score2) / 2.0;
    return avg - analyzer.baseline;
}

int process_numbers(int limit) {
    int sum_val = 0;
    for (int i = 1; i <= limit; ++i) {
        sum_val += i;
    }
    return sum_val;
}

int main() {
    std::cout << "=== C++ NATIVE (-O3) ===" << std::endl;
    double delta = compute_performance("Core", 50.0, 85.0, 95.0);
    int total_sum = process_numbers(100);
    std::cout << "Delta: " << delta << std::endl;
    std::cout << "Sum: " << total_sum << std::endl;
    return 1;
}


🔍 VII. IN-DEPTH SYNTAX, ADVANTAGES & LIMITATIONS ANALYSIS

⚡ 1. TokenVector (.tkv)

🐍 2. Python 3 (.py)

⚡ 3. C++20 (.cpp)


🛠️ VIII. OPERATIONAL & BUILD GUIDE

1. Compiling a .tkv Source File via tkvc.exe:

.\release\3.code\dist\tkvc.exe build release\3.code\examples\e2e_test.tkv

2. Executing the Compiled Native .exe:

.\release\3.code\examples\e2e_test.exe


🔗 IX. .NET ECOSYSTEM INTEROPERABILITY (.NET INTEROP)

Beyond compiling native Python-syntax code ahead-of-time, TokenVector provides seamless direct invocation of any .NET/NuGet library without requiring manual C wrappers:

Production Verification: RamGuard — an automated background RAM monitoring and working-set trimming service for Windows, re-engineered 100% in .tkv (zero Python code remaining). It leverages Process and ComputerInfo from the .NET BCL via __tkv_extern_class__, validated end-to-end (logging, cooldown loops, structured try/except error handling) — an operational utility, not a mock demo (private proprietary project).


📌 X. 3-POLE BENCHMARK SUMMARY

  1. CPython 3.12: Ideal for rapid automation scripting, exploratory prototyping, and data science research. Trade-offs: Lower execution speed, bloated distribution artifacts (tens of MBs), and severe multicore limitations due to GIL contention.
  2. TokenVector AOT: The ideal bridge between both worlds! Retains 100% of Python’s developer ergonomics while generating ultra-compact .exe artifacts (tens of KBs), delivering ~25.9× FASTER multithreaded integer execution (empirically measured) by eliminating the GIL, backed by native support for yield from, async/await, ctypes FFI, and first-class .NET ecosystem interop.
  3. C++ Native: Delivers unconstrained raw computational throughput and deterministic manual memory governance, but demands high syntactical friction, extended build cycles (multi-second toolchain latency), and significantly elevated development costs.

```