Large Language Models (LLMs) like GPT-4 or Llama 3 are world-famous for writing textual data, but underneath the surface, they text is represented as tokens i.e. numerial arrays. If you map geometric shapes and spatial coordinate layouts into sequence-based text strings, a raw Transformer can easily learn how to design structural architectural blueprints.

In this guide, you will build a custom Generative Transformer architecture from scratch using PyTorch to act as an automated interior designer. Given a text prompt size limitation like [50x30], the model automatically calculates and places bedrooms, kitchens, bathrooms, and living rooms without clipping windows or overlapping walls. Since you are building it from scratch locally, it won’t be as fancy as other models in the market, but it will provide overall understanding.

Below is the deep, step-by-step engineering breakdown of how this system evolved, the real-world bugs encountered during execution, and how you resolved them using sound machine learning design patterns.


💡 Why Build This When You Can Just Ask ChatGPT?

Sure, you can ask an AI to write a Python script that renders a house blueprint. But copy-pasting code won’t teach you how to architect a neural network.

This post isn’t just a code dump. It is a deep engineering case study on how to train, troubleshoot, and optimize a raw Transformer model on your local machine. You will explore why industry-standard tools like OpenAI’s tiktoken might not be good alignment, how to diagnose recursive token loops, and how to structure sparse spatial data to a tiny model on local hardware.


Phase 1: The Tokenizer Selection (Why Tiktoken Fails at Math)

Computers doesn’t understand text, rather the text is encoded into numbers. This process of converting text to numbers is called as Tokenization, whose job is to split strings into tokens and map them to numerical integer IDs.

For any text application, tiktoken (OpenAI’s Byte Pair Encoding tokenizer) is the standard library to use. However in our case it’s a misfit, as our input isnt text but spatial coordinates.

The Tokenizer Selection (The VRAM & Compute Bottleneck)

Before writing any neural network layers, you have to solve the fundamental data problem: Tokenization. While a live terminal test reveals that OpenAI’s standard GPT-2 tokenizer ( tiktoken) cleanly groups complete digit sequences like "185" without fragmentation, using it for a scratch-built domain model creates a massive computational bottleneck: Vocabulary Bloat.

The GPT-2 tokenizer forces a fixed dictionary size of 50,257 unique tokens. In a PyTorch Transformer, the final linear projection head (nn.Linear(embed_dim, vocab_size)) must calculate a probability score across the entire vocabulary at every single sequence step. This eats up VRAM and drastically slows down training on consumer hardware.

For an embedding size of 128, using tiktoken explodes this final matrix to 128 × 50,257—forcing a tiny model to waste valuable GPU parameters tracking millions of unrelated internet text words. By building a custom word-level tokenizer, you restrict the vocabulary size strictly to ~50 active layout targets, ensuring the model trains efficiently on local consumer hardware.

The Solution: A Custom Syntax Tokenizer

By building a dedicated regex tokenizer, you limit the model’s vocabulary strictly to the structural words, numbers, and symbols used in our layouts. This drops our vocabulary size from 100,000 to around 60 unique tokens, forcing the network to train much faster with sharp, mathematical precision.


Phase 2: The Initial Code & The Infinite Number Loop Regression

With our token dictionary optimized, you built a character-level data pipeline inside dataset_and_tokenizer.py and connected it to our Causal Transformer model.

The Original Text-Level Data Pipeline (dataset_and_tokenizer.py)

import random
import os

def generate_single_house_string():
    canvas_w, canvas_h = 400, 400
    split_y = random.randint(180, 220)
    top_split_x = random.randint(160, 240)
    bot_split_x = random.randint(160, 240)
    
    colors = {
        "living": ["#a2d2ff", "#bde0fe"],
        "bed":    ["#ffafcc", "#ffc8dd"],
        "kitchen":["#b7e4c7", "#d8f3dc"],
        "bath":   ["#fcd5ce", "#ffe5d9"]
    }
    
    top_rooms = ["living", "bed"]
    random.shuffle(top_rooms)
    bot_rooms = ["kitchen", "bath"]
    random.shuffle(bot_rooms)
    
    layout = [
        {"type": top_rooms, "x": 20, "y": 20, "w": top_split_x - 20, "h": split_y - 20, "color": random.choice(colors[top_rooms])},
        {"type": top_rooms, "x": top_split_x, "y": 20, "w": canvas_w - top_split_x - 20, "h": split_y - 20, "color": random.choice(colors[top_rooms])},
        {"type": bot_rooms, "x": 20, "y": split_y, "w": bot_split_x - 20, "h": canvas_h - split_y - 20, "color": random.choice(colors[bot_rooms])},
        {"type": bot_rooms, "x": bot_split_x, "y": split_y, "w": canvas_w - bot_split_x - 20, "h": canvas_h - split_y - 20, "color": random.choice(colors[bot_rooms])}
    ]
    
    svg = f'<svg width="{canvas_w}" height="{canvas_h}" xmlns="http://w3.org"><house>'
    for r in layout:
        name = r["type"].upper()[:4]
        svg += f'<rect x="{r["x"]}" y="{r["y"]}" w="{r["w"]}" h="{r["h"]}" fill="{r["color"]}" stroke="#333" stroke-w="3"/>'
        svg += f'<text x="{r["x"] + 15}" y="{r["y"] + 30}" font-s="14">{name}</text>'
    svg += '</house></svg>'
    return svg

The Bug: Infinite Token Tailspins

When executing inference using this format, the model generated the initial layers correctly but suddenly devolved into an infinite character loop:

<rect x="20" y="20" w="168" h="174" fill="#ffc8dd" stroke="#333" stroke-w="#a"#at st st sttroke="185550"#a20"145oke5

What Changed and Why?

• The Culprit: The character-level tokenizer. A character tokenizer processes “184” as three separate sequential items: ‘1’, ‘8’, and ‘7’. Because numbers look completely identical across different rooms in the training file, the Causal Attention Layer lost its orientation. It fell into a state loop, continually guessing numbers because it could not track string context across deep layouts. • The Fix: Swapping the raw character layout for a Syntax-Aware Tokenizer using regular expressions (re). This system bundles entire numbers (like “184”), distinct tags (like <rect), and color hex codes into single individual word tokens. This reduced the sequence layout context length from 350 to under 90 tokens, making tracking infinitely easier for the neural layers.


Phase 3: Structural Degradation (The Context Deficit)

With whole numbers being parsed as singular, clean units, the raw looping issue vanished. However, a new challenge cropped up. The network successfully drew the living room, bedroom, and bathroom, but broke down on the final room block:

<text x="228" y="221" font-s= 14 > BED </text >
<rect x="20" y="215" w="179" h="165" fill="#fcd5ce" stroke="#333" ...
<text x="214" y="245" font-s= 14 > KITC w="150" stroke="#333" stroke-w= 14 > KITC

The Bug: Attribute Scrambling from Verbose Markup

The model scrambled its attributes together (KITC w=”150” stroke=). Because our custom transformer model was quite small (4 layers, 128 embedding size), forcing it to predict non-functional HTML structures like stroke-width=”3” over and over caused context fatigue. The attention matrix lost track of whether it was rendering an object bounding rectangle or a text coordinate label.

When you ask a model to predict raw SVG/HTML strings, you aren’t just asking it to perform spatial math (calculating coordinates like x, y, w, h). Rather, you are forcing it to memorize a highly verbose, repetitive linguistic template. A small model (4 layers, 128 embedding size) has a limited “cognitive capacity.” It can only keep track of a few patterns at once. When you force it to output endless lines of boilerplate code like stroke="#333" stroke-width="3", you are flooding its short-term memory window with repetitive text data that has zero functional value to the actual layout of the house.

What Changed and Why?

• The Culprit: Mixing core architectural data with verbose document decoration markup. • The Fix: Stripped the HTML syntax completely out of the neural network’s training environment i.e training the LLM to output a highly dense sequence of only raw spatial parameters:LIVI 20 20 181 195 BED 201 20 179 195 ... [END] Once the model predicts this dense token stream, a simple helper function takes the coordinates and programmatically wraps them into compliant, browser-ready SVG tags.


Phase 4: The Missing Link (The Empty Canvas Error)

After optimizing the text arrays down to crisp, clean line arrays, running the newly compiled training model threw a hard Python traceback crash:

Traceback (most recent call last):
  File "train_blueprint_llm.py", line 92, in <module>
    print(f"Epoch {epoch+1:02d} | Loss: {total_loss / len(dataloader):.4f}")
ZeroDivisionError: division by zero

The Bug: Dataset File Handling Desync

The file parsing parameters inside train_blueprint_llm.py were still configured to split dataset samples using an old legacy delimiter (===END_HOUSE===). The fix eliminated all the bulky HTML formatting and packed each house layout into a single, dense, self-contained row of numbers and labels and stopped generating the ===END_HOUSE=== string altogether.

When the old code loader ran on the new dataset file, it searched for ===END_HOUSE===, found zero matches, and wrapped the entire dataset file into an empty list. When PyTorch initialized the DataLoader mini-batch manager, it calculated the number of operational steps by dividing the data length by the batch size. Because the list length was 0, the engine hit a fatal mathematical paradox: ZeroDivisionError: division by zero, crashing the terminal before training could even start.

What Changed and Why?

• The Fix: Rewrote the HouseBlueprintDataset loader to handle samples line-by-line (line.strip()), fully syncing it with the streamlined structure generator. So, instead of loading the entire hard drive file into memory and searching for string tags, the updated script scans the text file sequentially row-by-row. The line.strip() helper instantly strips out hidden white space characters and trailing system linebreaks (\n), while the conditional if line.strip() verifies that the model completely skips over empty rows or accidental trailing blank lines.


The Complete, Working Production Pipeline

Below is the final synchronized, fully functional production pipeline. It is split into three clean files that you can run locally on any standard computer setup.

1. Data Synthesis & Tokenizer Engine (dataset_and_tokenizer.py)

Run this script first. It generates a completely optimized text file called house_dataset.txt containing 2,000 dense, mathematically stable floor plans.

import random
import os

def generate_single_house_string():
    """Generates a highly dense, token-friendly layout string for stable attention tracking."""
    split_y = random.randint(180, 220)
    top_split_x = random.randint(160, 240)
    bot_split_x = random.randint(160, 240)
    
    # Structural configuration mappings
    colors = {"LIVI": "#a2d2ff", "BED": "#ffc8dd", "KITC": "#b7e4c7", "BATH": "#fcd5ce"}
    
    layout = [
        {"type": "LIVI", "x": 20, "y": 20, "w": top_split_x - 20, "h": split_y - 20},
        {"type": "BED",  "x": top_split_x, "y": 20, "w": 400 - top_split_x - 20, "h": split_y - 20},
        {"type": "KITC", "x": 20, "y": split_y, "w": bot_split_x - 20, "h": 400 - split_y - 20},
        {"type": "BATH", "x": bot_split_x, "y": split_y, "w": 400 - bot_split_x - 20, "h": 400 - split_y - 20}
    ]
    
    svg_tokens = ["[50x30]"] # The size constraint target prefix prompt
    for r in layout:
        svg_tokens.append(f"{r['type']} {r['x']} {r['y']} {r['w']} {r['h']}")
    svg_tokens.append("[END]")
    
    return " ".join(svg_tokens)

class BlueprintTokenizer:
    def __init__(self, dataset_path="house_dataset.txt"):
        if not os.path.exists(dataset_path):
            self.words = ['[PAD]', '[50x30]', 'LIVI', 'BED', 'KITC', 'BATH', '[END]']
        else:
            with open(dataset_path, "r") as f:
                text = f.read()
            self.words = sorted(list(set(text.split())))
            if '[PAD]' not in self.words: self.words.append('[PAD]')
            
        self.vocab_size = len(self.words)
        self.word_to_id = {w: i for i, w in enumerate(self.words)}
        self.id_to_word = {i: w for i, w in enumerate(self.words)}
        self.pad_id = self.word_to_id['[PAD]']

    def encode(self, text_string):
        return [self.word_to_id[w] for w in text_string.split() if w in self.word_to_id]

    def decode(self, id_list):
        return " ".join([self.id_to_word.get(i, '') for i in id_list if self.id_to_word.get(i, '') != '[PAD]'])

if __name__ == "__main__":
    print("Generating 2,000 anchored house layout samples...")
    with open("house_dataset.txt", "w") as f:
        for _ in range(2000):
            f.write(generate_single_house_string() + "\n")
    print("Dataset successfully created as 'house_dataset.txt'!")

Code Breakdown: dataset_and_tokenizer.py

• generate_single_house_string(): Sets a fixed bounding box of 400 × 400 pixels. Slicing it randomly ensures every sample has slightly different layout sizes, forcing the LLM to learn mathematical boundary rules rather than memorizing static digits. The function loops through the calculated coordinates and wraps them into a single string separated by spaces. • BlueprintTokenizer: Reads the output file and tracks uniquely observed string units. Instead of breaking “184” into ‘1’, ‘8’, and ‘4’ like a legacy character encoder (which causes infinite loop regressions during inference), it maps “184” to a unique index entry. This reduces sequence complexity down to just 25 active tokens per house blueprint.

2. Core Transformer Deep Learning Loop (train_blueprint_llm.py)

Run this second. It reads the dense token dataset line-by-line, trains the network over 15 epochs using an AdamW optimizer, and outputs the final weight parameters as a serialization file.

import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import Dataset, DataLoader
import os
import warnings

# Clear terminal warning print overheads
warnings.filterwarnings("ignore", category=UserWarning, module="torch.nn.modules.transformer")
from dataset_and_tokenizer import BlueprintTokenizer

class HouseBlueprintDataset(Dataset):
    def __init__(self, dataset_path, tokenizer, max_length=25):
        self.tokenizer = tokenizer
        self.max_length = max_length
        with open(dataset_path, "r") as f:
            self.samples = [line.strip() for line in f if line.strip()]

    def __len__(self): return len(self.samples)

    def __getitem__(self, idx):
        encoded = self.tokenizer.encode(self.samples[idx])
        if len(encoded) > self.max_length:
            encoded = encoded[:self.max_length]
        else:
            encoded = encoded + [self.tokenizer.pad_id] * (self.max_length - len(encoded))
        tensor_data = torch.tensor(encoded, dtype=torch.long)
        return tensor_data[:-1], tensor_data[1:]

class BlueprintTransformer(nn.Module):
    def __init__(self, vocab_size, embed_dim=128, num_heads=4, num_layers=4, max_seq_len=25):
        super().__init__()
        self.token_embeddings = nn.Embedding(vocab_size, embed_dim)
        self.position_embeddings = nn.Embedding(max_seq_len, embed_dim)
        
        encoder_layer = nn.TransformerEncoderLayer(
            d_model=embed_dim, nhead=num_heads, dim_feedforward=embed_dim * 4, 
            dropout=0.1, activation='gelu', batch_first=True, norm_first=True
        )
        self.transformer_blocks = nn.TransformerEncoder(encoder_layer, num_layers=num_layers)
        self.ln_f = nn.LayerNorm(embed_dim)
        self.lm_head = nn.Linear(embed_dim, vocab_size)

    def forward(self, x):
        batch_size, seq_len = x.size()
        device = x.device
        pos = torch.arange(0, seq_len, dtype=torch.long, device=device).unsqueeze(0)
        x = self.token_embeddings(x) + self.position_embeddings(pos)
        
        mask = torch.triu(torch.ones(seq_len, seq_len, device=device), diagonal=1).masked_fill(mask == 1, float('-inf'))
        mask = torch.triu(torch.ones(seq_len, seq_len, device=device), diagonal=1).masked_fill(mask == 1, float('-inf'))
        
        x = self.transformer_blocks(x, mask=mask, is_causal=True)
        return self.lm_head(self.ln_f(x))

if __name__ == "__main__":
    BATCH_SIZE, EPOCHS, LEARNING_RATE, MAX_LEN = 32, 15, 5e-4, 25
    DEVICE = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    
    tokenizer = BlueprintTokenizer("house_dataset.txt")
    dataset = HouseBlueprintDataset("house_dataset.txt", tokenizer, max_length=MAX_LEN)
    dataloader = DataLoader(dataset, batch_size=BATCH_SIZE, shuffle=True, drop_last=False)
    
    model = BlueprintTransformer(vocab_size=tokenizer.vocab_size, max_seq_len=MAX_LEN).to(DEVICE)
    criterion = nn.CrossEntropyLoss(ignore_index=tokenizer.pad_id)
    optimizer = optim.AdamW(model.parameters(), lr=LEARNING_RATE, weight_decay=0.01)
    
    print(f"Training Blueprint LLM on {DEVICE}...")
    model.train()
    for epoch in range(EPOCHS):
        total_loss = 0
        for batch_x, batch_y in dataloader:
            batch_x, batch_y = batch_x.to(DEVICE), batch_y.to(DEVICE)
            optimizer.zero_grad()
            loss = criterion(model(batch_x).view(-1, tokenizer.vocab_size), batch_y.view(-1))
            loss.backward()
            torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
            optimizer.step()
            total_loss += loss.item()
        print(f"Epoch {epoch+1:02d}/{EPOCHS:02d} | Cross-Entropy Loss: {total_loss / len(dataloader):.4f}")
        
    torch.save(model.state_dict(), "blueprint_llm_weights.pth")
    print("Training Complete! Weight matrix saved to 'blueprint_llm_weights.pth'")

Code Breakdown: train_blueprint_llm.py

• HouseBlueprintDataset: Reads tokens line-by-line. If a line is shorter than MAX_LEN (25 tokens), it fills the trailing items with the padding index. Autoregressive sequence shifting forces the model to learn context dependencies: given tokens from 0 to N-1, it must accurately predict the target token array from 1 to N. • BlueprintTransformer: Maps incoming array coordinates through the embedding matrices. A strict Causal Mask blocks token positions from reading forward in time during multi-head attention evaluations. Setting norm_first=True ensures smooth parameter gradient changes throughout deep layer adjustments, preventing training loss values from exploding into NaN figures.

3. Autoregressive Generator & Web Compiler (generate_house.py)

Run this last. It passes your input seed prompt constraint ([50x30]) directly into the trained model weights, decodes the spatial parameter array autoregressively, and writes it directly to a clean browser-compliant HTML file.

import torch
import torch.nn.functional as F
import sys
from dataset_and_tokenizer import BlueprintTokenizer
from train_blueprint_llm import BlueprintTransformer

def compile_tokens_to_svg(token_string):
    """Compiles clean structural data coordinates directly into a browser-valid HTML vector layout."""
    colors = {"LIVI": "#a2d2ff", "BED": "#ffc8dd", "KITC": "#b7e4c7", "BATH": "#fcd5ce"}
    html_output = '<svg width="400" height="400" xmlns="http://w3.org">\n'
    
    words = token_string.split()
    for i in range(1, len(words) - 4, 5):
        try:
            room_type = words[i]
            x, y, w, h = words[i+1], words[i+2], words[i+3], words[i+4]
            color = colors.get(room_type, "#ffffff")
            
            html_output += f'  <rect x="{x}" y="{y}" width="{w}" height="{h}" fill="{color}" stroke="#333" stroke-width="3"/>\n'
            html_output += f'  <text x="{int(x)+15}" y="{int(y)+30}" font-family="sans-serif" font-size="14" font-weight="bold" fill="#333">{room_type}</text>\n'
        except (ValueError, IndexError):
            continue 
            
    html_output += "</svg>"
    return html_output

def generate_house_layout(prompt_text):
    DEVICE = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    tokenizer = BlueprintTokenizer("house_dataset.txt")
    
    model = BlueprintTransformer(vocab_size=tokenizer.vocab_size, max_seq_len=25).to(DEVICE)
    model.load_state_dict(torch.load("blueprint_llm_weights.pth", map_location=DEVICE))
    model.eval()
    
    input_ids = tokenizer.encode(prompt_text)
    input_tensor = torch.tensor([input_ids], dtype=torch.long, device=DEVICE)
    generated_ids = input_ids.copy()
    
    print("\n--- Autoregressive Decoding Data Stream ---")
    with torch.no_grad():
        for _ in range(25):
            logits = model(input_tensor)
            next_token_logits = logits[:, -1, :] / 0.6
            
            # Simple Top-K sorting to clear noise
            v, _ = torch.topk(next_token_logits, 2)
            next_token_logits[next_token_logits < v[:, [-1]]] = float('-inf')
            
            next_token_id = torch.multinomial(F.softmax(next_token_logits, dim=-1), num_samples=1).item()
            generated_ids.append(next_token_id)
            
            new_word = tokenizer.id_to_word.get(next_token_id, '')
            sys.stdout.write(f"{new_word} ")
            sys.stdout.flush()
            
            if new_word == "[END]": break
            input_tensor = torch.cat([input_tensor, torch.tensor([[next_token_id]], dtype=torch.long, device=DEVICE)], dim=1)
            
    with open("ai_generated_layout.html", "w") as f:
        f.write(compile_tokens_to_svg(tokenizer.decode(generated_ids)))
    print("\n\n[Success] Layout rendered and stored to: 'ai_generated_layout.html'")

if __name__ == "__main__":
    generate_house_layout(prompt_text="[50x30]")
    

Code Breakdown: generate_house.py

• generate_house_layout(): Initialise the network framework and loads the raw serialized weights. The generation loop takes the prompt tensor, extracts the predictions at the absolute final sequence position (logits[:, -1, :]), applies a Top-K filter to wipe out illogical options, and samples from a clean Softmax probability distribution. • compile_tokens_to_svg(): Parses the clean predicted layout tokens. By stepping through the string array in structured jumps of 5 (Room_Type → X → Y → Width → Height), it reads raw semantic coordinates and structures them into standard, browser-compliant XML objects.

Conclusion: The Engineering Takeaway

Building this project highlights why vertical domain token filtering is so critical in practical model design. Forcing a generalized text model to manage complex structural syntax variations or verbose document decoration tags wastes parameter capacity. By applying Linear Structure Anchoring—separating structural prediction from visual rendering—we reduced context window bloat and eliminated syntax loops entirely, turning a failing project into a robust custom application.