The ideas behind a count you can defend
Six ideas from the machine-learning textbook, translated for estimators. No equations. What the AI actually does, and what stays with you.
Six ideas from the machine-learning textbook, translated for estimators. No equations. What the AI actually does, and what stays with you.
If you are going to put an AI count in a bid, you deserve to know what kind of machine you are trusting. Not the marketing version, and not the maths either. The version in between: the handful of ideas that every symbol-detection tool draws on, and what each one means for the number you sign your name to.
The ideas below come from a standard university course in statistical machine learning, Andrew Moore's tutorial series at Carnegie Mellon, which has been online and freely used by teachers since 2001. We link the original slides at the foot of the page. What follows is our translation for the estimating desk.
One honest note before we start. This page explains the ideas any AI takeoff tool is built on. It is not a description of SimplyCount's internal engine. Everything we say about the product here repeats what we already commit to on the benchmark and trust pages: counting runs on your machine, you review every detection, and an unverified count is shown as a floor, never a total.
The oldest form of machine learning is also the most robust. You do not write rules describing a duplex receptacle. You show the machine one, and it finds everything on the sheet that looks most like it. The textbook calls this instance-based or memory-based learning, and it is over a century old because it works.
The consequence matters for construction. No two firms draw a legend the same way, and a fixed catalogue of "standard" symbols fails the moment a designer adds a tick mark. A tool that learns from your example instead of a vendor's rulebook copes with your legend, not an idealised one.
For your bid: teach a symbol once from your own legend and it is saved for every future job. Your library, and your speed, compound the more you count.
A classifier never says "this is a receptacle." It says "this is probably a receptacle, and here is how sure I am." Every mark on a counted sheet has a confidence behind it, whether or not the vendor shows it to you. The maths cannot produce certainty from a drawing, and any tool that claims it has hidden the uncertainty rather than removed it.
This is the single most useful thing to know when you evaluate a takeoff tool. The question is not "how accurate is it," which is a number no one can defend in a bid. The question is "what does it do with the detections it is not sure about."
For your bid: a count that is not fully verified is shown as a floor (≥), never dressed up as a total. The floor is the maths being honest with you.
The textbook's example is a robot with unreliable proximity sensors. The true state of the world is hidden. All the robot gets is a noisy signal of it, and the job of the model is to work out the most probable truth from that signal.
A plan set is exactly this problem. The drawing the engineer intended is the hidden truth. What arrives on your desk is a scan with faded lines, symbols overlapping text, a legend photocopied three times. A counting tool is reasoning about what is most probably there, and that reasoning is where mistakes hide.
For your bid: your review is the last and best sensor. You see the marked sheet, accept or reject each detection, and the total that leaves the app is the one you approved.
A model that has memorised its examples looks perfect on them and fails on anything new. The textbook calls this overfitting, and the whole discipline of cross-validation exists because the only honest test of a model is data it has never seen. A vendor's own demo sheet is, by definition, data the vendor has seen.
This turns a technical idea into a purchasing rule. Never judge a takeoff tool on the plan it was tuned on. Judge it on yours.
For your bid: our benchmark is published with its assumption stated, and the strongest test is still your own plan set. Say so in your access request and the team sends back a marked-up count of your sheet.
Reinforcement learning is the study of a controller that must, in the textbook's words, suck up some short-term punishment to reach a long-term reward. Two ideas carry over. Each decision produces feedback that improves the next one. And the value of an action is not what it pays today but what it pays over every job that follows, discounted a little for distance.
The first review pass on a new legend is the short-term cost. It is also the feedback. Every symbol you teach and every detection you confirm is banked for the next plan set, and the return on that afternoon is paid on every bid after it.
For your bid: SimplyCount is self-learning. Teach a symbol once and it counts that symbol on every future job, getting more accurate the more you use it.
The textbook's first classifier asks one question at a time and picks the question that splits the examples most cleanly. It calls the payoff of a good question information gain. The lesson underneath it is that a class is only as learnable as the features that separate it from its neighbours.
Two symbols that differ by a single small tick, or a symbol that changes meaning with a subscript, carry very little information gain. They are hard for a machine for the same reason they are hard for a tired estimator at 11pm. Inconsistent symbols and standards are what make hand counts hard to defend, and the machine does not escape that difficulty. It just declares it.
For your bid: teach both variants as separate symbols, review the ones that look alike, and let the tool flag rather than guess. A defensible count is one where the hard cases were decided by you.
Each one comes straight from an idea above. A straight answer is a good sign. A percentage is not.
Idea 2. If the answer is "it counts it," the uncertainty went into your bid. Look for a floor, a flag, or a review queue.
Idea 4. The vendor's demo sheet is the training set. Yours is the exam.
Idea 1. A catalogue fails on the first non-standard symbol. Learning from your example does not.
Ideas 3 and 5. You should be able to reject a detection on the sheet, and the correction should carry to the next job.
Not a textbook idea, but the one IT will ask first. SimplyCount counts on your own machine. How your plans stay yours →
| Idea | What it means for your bid | Where you see it |
|---|---|---|
| Instance-based learning | Learns your legend from one example, not a vendor catalogue | Teach a symbol once; saved for every future job |
| Probability, Bayes classifiers | Every mark carries a confidence; honesty means showing it | Unverified counts shown as a floor (≥), never a total |
| Hidden Markov models | A scan is a noisy signal of the intended drawing | You review each detection on the marked sheet and sign off |
| Cross-validation | Only unseen plans are a real test | Published benchmark with its assumption; pilot on your own plans |
| Reinforcement learning | Today's review is tomorrow's speed | Your reviews build your legend library, and that library is what speeds the next count |
| Decision trees, information gain | Look-alike symbols are hard for everyone | Teach variants separately; the tool flags rather than guesses |
The left column is the textbook. The right column is what the product page already promises. Nothing in between is a claim about internals.
Every idea above is taken from one lecture in Andrew Moore's Statistical Data Mining Tutorials at Carnegie Mellon University. The slides are free for anyone to read. If a member of your team wants the real thing, these are the six to start with, in the order we used them.
Terms you meet along the way are defined for estimators in our glossary.
Request early access with your work email, attach a sample sheet if you want a pilot count, and the team sends back a marked-up sheet you can review yourself. 168 pages · 60 symbols · ~1 hour.
Free during early access · No credit card · Counting runs on your machine