01 / SYSTEM OVERVIEWA self-evolving AI network.
Clawsyn is building a domain-focused language model and agent network for crypto, finance and code. The Progenitor model is the shared foundation: crawlers expand its source corpus, training releases develop its capabilities, and specialized replicas will apply those capabilities to individual domains.
The development cycle connects data acquisition, corpus engineering, model optimization, evaluation and replica incubation. Protocol documentation, market data, onchain records and software repositories form the target knowledge domains.
02 / MODEL FOUNDATIONCS-1: a domain-native language model.
The current CS-1 checkpoint follows the team's Crawlnet model lineage and uses a nanochat-based GPT architecture. Its compact BF16 checkpoint provides a concrete foundation for domain adaptation, compression experiments and specialized deployment.
- PARAMETERS
- 122.6M
- TRANSFORMER LAYERS
- 10
- HIDDEN DIMENSION
- 640
- ATTENTION HEADS
- 5
- CONTEXT WINDOW
- 2,048 tokens
- BPE VOCABULARY
- 16,384 tokens
Architecture values come from the current model configuration and weight tensors. Training statistics below are recorded in the checkpoint's model card and training summary.
03 / TRAINING PIPELINEDomain data, grounded supervision.
Corpus-native tokenization
A byte-pair encoding tokenizer trained on the domain corpus defines a 16,384-token vocabulary for the model's text representation. The recorded pretraining corpus combines crawled crypto pages with public market, funding and onchain records serialized as text.
Domain-native pretraining
The model card records training from scratch on 134,994 cleaned pages: 89,367 crawler pages and 45,627 record pages. It reports 191.5M unique training tokens and 626.0M processed tokens across approximately 3.27 passes through the corpus.
Grounded instruction tuning
The supervised fine-tuning pipeline combines synthetic question-answer examples grounded in source pages, extractive examples derived from page headings, and open-book conversations with source context. This provides both domain-language exposure and supervision for answering from supplied evidence.
Versioned checkpoint provenance
The release package includes model configuration, tokenizer artifacts, dataset metadata, training summaries and file checksums. These artifacts preserve the checkpoint's architecture and training lineage for subsequent Clawsyn branches.
04 / CONTINUAL LEARNING ROADMAPControlled model evolution.
The next-stage architecture separates retrieval from weight updates. Retrieval-augmented generation (RAG) will supply relevant source context at inference time, while continual pretraining will incorporate curated new data into scheduled model candidates.
The planned training loop adds semantic deduplication, domain balancing and historical data replay. Held-out tasks, regression checks and checkpoint comparison will determine whether a candidate can become the next Progenitor release. Crawled data enters the corpus first; model upgrades follow a separate training and evaluation process.
Parameter-efficient fine-tuning (PEFT), including low-rank adaptation (LoRA), is the proposed specialization path. Structured pruning and distillation are additional compression research directions, subject to capability and runtime evaluation.
These modules are development targets. The current interface does not run autonomous retraining or promote checkpoints.
05 / REPLICA ARCHITECTUREShared foundation. Specialized execution.
A planned replica combines a versioned Progenitor checkpoint with a domain adapter, retrieval memory, task policy and allocated compute budget. An incubation event initializes that configuration and its ownership record.
01 / RESEARCHProtocol intelligence
Protocol documentation, onchain activity and digital-asset research, including NFT analysis.
02 / FINANCEQuantitative analysis
Market research, strategy evaluation and backtesting. Execution integrations will require explicit permissions.
03 / ENGINEERINGSoftware development
Repository analysis, code generation and testing through controlled development tools.
Specialization, tool integration and task evaluations will define each replica's capabilities. Validated task outcomes can become candidates for future training data and Progenitor upgrades.
06 / DEVELOPMENT STATUSFrom model foundation to network.
A CS-1 text-model checkpoint is available within the project. This website currently provides captured-source inspection, sample wallet records and wallet connection. The Progenitor blueprint illustrates the planned lifecycle. Its scheduled progress display is illustrative and does not measure training completion.
Inference integration, automated training orchestration, replica deployment and mint settlement are subsequent milestones. Incubation capacity, ownership rights, operating costs and token allocations will be specified before minting opens.